A conjecture proposed in 1939 fell on July 20, disproven by a counterexample short enough to fit in a single social-media post, from a mathematician working with Claude Fable 5. Twelve days later, OpenAI answered with volume: ten solved problems, each with a machine-checkable proof attached, for a total compute cost of roughly $2,000. Weeks earlier, on June 2, the International Mathematical Union had endorsed the Leiden Declaration on Artificial Intelligence and Mathematics, an attempt to write rules for a pace nobody had set them for yet.
The Thesis
Mathematics is working through a problem that reaches every institution certifying expert output: what kind of review a result must go through before it is accepted, what evidence a student must produce before a credential is issued, and who controls the verification system behind both. The field produced answers to all three in a few months, written by mathematicians under pressure.
The Signal
Three developments worth watching this week.
What happened. In July 2025, several AI models solved five of six problems at the International Mathematical Olympiad, and Terence Tao, one of the world’s leading mathematicians, described 2025 as the year the tools became broadly useful. In October 2025, Ernest Ryu, then a mathematician at UCLA, spent roughly 12 hours with ChatGPT and solved a problem that had been open for 42 years. Tao, Javier Gómez-Serrano of Brown, and two DeepMind mathematicians ran Google’s AlphaEvolve on 67 open problems, beating the best known answer on 23 and matching it on 36 in a November 2025 paper. In January 2026, Ravi Vakil, president of the American Mathematical Society, and colleagues proved a new result using two Gemini-based systems.
Why it matters. A year of AI development separates high-school competition problems from conjectures that had defeated professional mathematicians for decades. Gómez-Serrano estimates that roughly two-thirds of his working time now involves AI tools, and notes that a specialist might have obtained any one of the AlphaEvolve results in a few months, while his team obtained comparable results across unfamiliar fields in a day or two. Grant cycles, tenure review, and hiring committees are calibrated to a rate of output that is being exceeded.
Second-order effect. The rate of discovery and the publication rate are diverging. The July counterexample was public, checked, and analysed by a Fields medallist within 48 hours, with a preprint following and peer review still pending. An institution that measures research productivity by publication is behind its own field. This affects promotion decisions, national assessment exercises, and the allocation of competitive funding.
What happened. On July 20, Levent Alpöge, a number theorist who works at Anthropic, posted a counterexample to the Jacobian conjecture, crediting Claude Fable 5 for “working during the World Cup final.” Ott-Heinrich Keller proposed the conjecture in 1939, and T. T. Moh predicted in 2008 that a human solution might take another century. Other mathematicians checked it by hand and by machine within hours, and Tao published an explanation the following day. The work settles most of the conjecture and has not cleared journal peer review.
Twelve days later, OpenAI released ten results in mathematics and theoretical computer science produced by an internal version of Astra, its next major model, each problem open for at least a decade. OpenAI published a 249-page manuscript and a machine-checkable certificate for every result on GitHub under an open licence, at a compute cost Greg Brockman put at roughly $2,000. Eleven weeks earlier, when a model overturned a long-standing Erdős conjecture, validation required nine external mathematicians signing off.
Why it matters. Both results were accepted on the merits that a third party could check independently. Alpöge’s counterexample is short enough that any competent reader can confirm it in an evening with free software. Astra’s proofs are up to hundreds of pages, so OpenAI supplied certificates that a proof assistant can check line by line. The question for any other institution is what its own version of a machine-checkable certificate would be, and whether one exists. In law, medicine, engineering, and government, the answer is still that a reviewing authority reads the document and forms a judgment.
Second-order effect. The ability to put an argument in machine-checkable form becomes increasingly important. It remains a complex skill, unevenly distributed across institutions and countries. José Simental of the National Autonomous University of Mexico has said publicly that his access to AlphaEvolve came through a semester spent in the United States, and has urged the community to take equitable access seriously. OpenAI’s offer of free frontier access for 100,000 academic researchers through 2027 addresses that partially and routes a share of national research capability through one company.
What happened. On June 2, an international group published the Leiden Declaration on Artificial Intelligence and Mathematics, and it was endorsed by the International Mathematical Union on the same day. It addresses verification, attribution, disclosure, research autonomy, funding, and unequal access. The Society for Industrial and Applied Mathematics’s editorial policy on AI requires authors to be human and able to defend every aspect of a submission directly, assigns responsibility for correctness to the authors, and bars reviewers from entering manuscripts into generative AI tools on confidentiality grounds. Thomas Bloom’s Erdős Problems database publishes categories separating cases where a model did all the work from genuine collaboration and ordinary assistance.
The limits are visible too. Some mathematicians have documented problems reported as solved autonomously that had in fact been solved years earlier, and cautioned that visible successes sit alongside a massive volume of unreported failures. Igor Pak doubts these systems will touch the hardest problems for a very long time, while Joel David Hamkins describes journals being overwhelmed by low-quality AI-generated submissions. No Millennium Prize problem has been solved.
Teaching changed faster than research. Hamkins stopped assigning homework because a substantial share of submissions now arrive AI-written. Robert Talbert of Grand Valley State concluded that he cannot meaningfully grade anything completed outside the classroom and now counts only in-class work. Princeton’s economics department moved to in-person-only final exams after students petitioned for the change. Oral examinations are returning across US campuses. Cornell’s Chris Schaffer, who introduced an oral defence in his biomedical engineering course, puts the rationale simply: “You won’t be able to AI your way through an oral exam.”
Why it matters. None waited for legislation. A professional body set norms, a publisher set submission conditions, and individual departments changed how credentials are earned. Each part is available to copy. Take-home work stopped functioning as evidence of learning for the same reason a plausible proof stopped functioning as evidence of a theorem, and both issues got the same answer: check in a setting where the output cannot be produced elsewhere.
Second-order effect. The training pipeline is the component institutions control most directly. Researchers acquire judgment by working through technical arguments as graduate students, and that work is what AI tools have the potential to absorb. Professors have warned that assigning problems a model solves instantly discourages students from building the underlying capacity. Ken Ono pairs optimism about research with concern about training. One researcher put the risk to Quanta anonymously: accelerating existing researchers whilst failing to produce new ones. Faculty departures compound it, with mathematicians moving to OpenAI, Google, and startups.
The Playbook
Four moves institutions can copy now.
The solutions are ordinary ones: in-class assessment, oral defences, and in-person finals. Decide which qualifications certify a capability a graduate must hold independently, and move assessment for those in person first. Princeton’s economics change came from a student petition, which suggests the demand may already exist internally.
The software libraries that check formal proofs, the tooling that translates arguments into them, and the expertise to use both are becoming the infrastructure by which AI-assisted research is confirmed. They are currently maintained largely by volunteer communities. Fund them as research infrastructure.
A commercial offer of free frontier access to 100,000 academic researchers is a public good and a risky dependency at the same time. Ask what happens at renewal, what happens to researchers outside the programme, and which national research capabilities would degrade if the terms changed.
The Leiden Declaration supplies norms; SIAM’s editorial policy supplies enforceable conditions. Each was written by practitioners under pressure, each is public, and adapting one costs a group a few weeks. Drafting an equivalent from scratch takes longer.
The Verification Test
“Our institution can verify the AI-assisted work it accepts and certifies.”
Test. Take recent items your institution put its name to: a funded grant report, a published paper, and a conferred qualification. For each, ask which part was independently confirmed, by whom, and which error class the check would have caught.
Pass criteria. Each item has a named, retrievable verification separable from the output itself: an executed proof or test suite, a re-run analysis, a cited source someone opened, or an assessment conducted in a setting the candidate could not outsource.
Fail smell. The answer is a person’s review, and that person read the material and found it convincing but is not fully certain. Plausibility is the property AI optimises for, and a fluent, internally consistent, well-formatted submission contains no information about whether it is correct.
The Metric
What it measures. The compute cost, at current rates, of the tokens used to generate all ten Astra solutions to problems that had each been open for a decade or more.
Why it matters now. The figure comes from OpenAI and covers only successful attempts. Even as a floor, it sits in comparison with a standard research grant and funding cycle, and that comparison matters to anyone allocating public research money. The effect on who produces findings is already visible: Quanta reported on August 3 that most new Erdős results have come from hobbyists and undergraduates using publicly available models.
The Lens — Horizon Search Institute
OpenAI published formal certificates with all ten Astra results, eleven weeks after an earlier result in the same lineage required nine mathematicians to read and endorse it. Mathematics had proof-checking tools available to adopt. Whether comparable tools exist for legal, clinical, or financial work is the open question, and the one worth tracking.
The cost is showing up first in training. Homework stopped functioning as evidence of learning, oral examinations returned, and mathematicians are concerned about the apprenticeship pipeline that produces the next generation of experts.
Institutional self-governance is one route to setting adoption terms on a shorter timeline than legislation. AI’s effects on the next generation of scientists and researchers — including who controls the models they depend on — will shape national power.
Links Worth Your Time
-
The Leiden Declaration on Artificial Intelligence and Mathematics
Short and written to be adapted. Any profession could rework its six concerns in an afternoon.
-
SIAM Publications: Editorial Policy on Artificial Intelligence
One page of enforceable submission conditions, including the requirement that authors be human and able to defend the work. The most directly copyable instrument in this issue.
-
The AI Revolution in Math Has Arrived
Quanta’s account of the arc from the 2025 Olympiad results forward, reported through the mathematicians themselves.
-
Perfect homework, blank stares
Associated Press on how quickly universities moved on assessment, and what they moved to.
-
Claude’s Cycles
Six pages from Donald Knuth, one of the field’s most exacting figures, describing his own reversal.
- OpenAI. Ten advances in mathematics and theoretical computer science. August 1, 2026; with 249-page manuscript and Lean 4 certificate repository (Apache 2.0).
- Kakaes, K. The AI Revolution in Math Has Arrived. Quanta Magazine, April 13, 2026.
- Why the Legendary Erdős Problems Are Falling to AI. Quanta Magazine, August 3, 2026.
- Alpöge, L. Counterexample to the Jacobian conjecture. Posted July 20, 2026; verification preprint, journal peer review pending.
- Tao, T. A digestion of the Jacobian conjecture counterexample. What’s New, July 21, 2026.
- Knuth, D. E. Claude’s Cycles. Stanford Computer Science Department, February 28, 2026; revised March 16, 2026.
- Leiden Declaration on Artificial Intelligence and Mathematics. June 2, 2026.
- International Mathematical Union. Adhering Organizations Circular Letter 8/2026. June 2, 2026.
- SIAM Publications. Editorial Policy on Artificial Intelligence.
- Jamnik, M. The Leiden Declaration: Mathematics, AI, and Making Our Values Explicit. Communications of the ACM, June 25, 2026.
- Tao, T. Post on AI tools and previously solved Erdős problems. Mathstodon.
- Bloom, T. F. Erdős Problems database and AI contributions taxonomy.
- Tao, T., Gómez-Serrano, J., Georgiev, A., & Wagner, F. Mathematical Exploration and Discovery at Scale. November 2025. arXiv:2511.02864.
- Ryu, E. Proof of convergence for Nesterov’s method. October 2025. arXiv:2510.23513.
- Litt, D. Mathematics in the Library of Babel. February 20, 2026.
- Gecker, J. Perfect homework, blank stares: Why colleges are turning to oral exams to combat AI. Associated Press, 2026.
- How AI Is Changing College Assessments of Proficiency. Government Technology, March 2026.
- The end of the take-home era: How exams have changed in the age of AI. The Daily Princetonian, February 25, 2026.
- Markman, J. OpenAI’s Astra Solved 10 Decades-Old Math Problems For Just $2,000. Forbes, August 3, 2026.