In the past year AI systems have produced genuine research results in mathematics. What is much less clear is how to measure what these systems can actually do, and how to use them in one's own work.
I will describe First Proof, which tests AI systems against unpublished mathematics, research mathematicians contribute problems with known but unpublished proofs, and the systems' solutions are refereed and graded, and report what successive batches of this have taught us.