News · June 11, 2026

Milestones workshop frames AI progress in mathematics

ICARM and Principia Labs convened mathematicians and AI researchers in San Francisco to define meaningful milestones for autonomous mathematics research.

In April, ICARM and Principia Labs co-sponsored the Milestones of Autonomous Mathematics Research workshop in San Francisco, bringing together mathematicians and AI researchers to ask how the field should assess increasingly capable mathematical AI systems. The workshop took place April 13-17, 2026, at Pebblebed Ventures and was organized by Elliot Glazer, Daniel Litt, Jacob Tsimerman, Alex Kontorovich, and Ravi Vakil.

The meeting responded to a rapidly changing landscape. Recent public discussion, including the New Scientist feature "A golden age of maths is dawning and mathematicians are freaking out," has emphasized both the excitement and anxiety around AI systems that can assist with, and in some cases contribute to, advanced mathematical work. Rather than treating isolated contest results or headline-making proofs as the only measures of progress, the workshop focused on what working mathematicians actually need from autonomous assistants and how those capabilities might be evaluated responsibly.

Participants worked toward a set of milestones across several dimensions of mathematical practice, including problem solving, generalization, exposition, formalization, and conjecturing. Later sessions included short reports on the state of the proposed milestones, discussion of the structure of test evaluations, and office hours for participants interested in contributing to the project.

A central theme was that AI progress in mathematics cannot be captured by a single benchmark. Useful evaluation will need to reflect the diversity of mathematical research, the role of human judgment in identifying important questions, and the need for clear standards that can help the broader community understand what current systems can and cannot do. Following the workshop, the organizers began work on a document intended to synthesize the discussion into credible benchmarks and milestones for the field.

← All news