News · June 11, 2026

FirstProof discussion finds a public home on ICARM Zulip

ICARM's public Zulip has become a working forum for mathematicians and AI researchers evaluating FirstProof claims, from Aletheia's results to Cursor, Archon, and other follow-on efforts.

The First Proof Project gave AI systems ten research-level mathematics problems and a short window to solve them. The result quickly became a test case for the field: news coverage emphasized both the ambition of the challenge and the mixed nature of the results, while mathematicians started asking the harder question of how such claims should be evaluated at all. ICARM's public #first proof Zulip channel has become one of the places where that evaluation is happening in public.

The most visible example is Aletheia, the mathematics research agent described by Tony Feng and collaborators. Their arXiv report states that Aletheia, powered by Gemini 3 Deep Think, solved six of the ten FirstProof problems according to majority expert assessments, with one problem not unanimously judged. The report links directly to ICARM Zulip threads for mathematical standards and problem-by-problem public comments, and Feng posted the Aletheia solutions in the channel as part of the external validation process.

That discussion did not stop with Aletheia. Others joined with their own attempts and follow-up work, including OpenAI solution discussions, Cursor's Problem 6 attempt by Shengtong Zhang and Wilson Lin, Archon's Lean formalizations of FirstProof Problems 4 and 6, Ingo Althofer and Dietmar Wolz's post-challenge "proof milking" experiments, Joseph Corneli's First Proof sprint, and newer work on proof verification and AI grading. The value of the channel is not that every claim is settled immediately; it is that claims are exposed to domain experts, competing interpretations, and a persistent public record.

The ICARM-hosted discussion has also helped sharpen what the next generation of AI-for-math benchmarks should measure. Participants debated whether a solution should meet an "accept with minor revisions" standard, how much human interaction is compatible with autonomy, how formal proof assistants such as Lean should be used, and how to avoid exhausting scarce expert referee time. Those questions are now part of the substance of FirstProof, not just commentary around it.

For ICARM, this is exactly the kind of infrastructure a computer-aided reasoning institute should provide: a neutral mathematical commons where fast-moving AI claims can be tested by people who understand the mathematics, the proof-assistant ecosystem, and the limits of current models. The public archive is already feeding reports, formalization work, and benchmark design. More importantly, it gives the mathematical community a place to turn headlines into shared knowledge.

Get involved in the public discussion at the ICARM Zulip: #first proof.

← All news