# When AI started doing real mathematics

> How far has AI got in research mathematics, and how far can its claims be trusted?

In about a year AI went from olympiad gold to refuting a famous conjecture and claiming a Millennium-problem case. Along the way, novelty checks, formal verification and credit disputes became as important as the results.

## Stops
1. **21 Jul 2025: Gemini Deep Think earns officially graded gold-medal score at IMO 2025** (Development) Certified olympiad-level proofs in natural language set the baseline. [page](https://wheresthe.ai/d/deepmind-gemini-imo-gold-2025-07/) · [source](https://deepmind.google/discover/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/)
2. **14 May 2025: Google DeepMind's AlphaEvolve agent improves algorithms and open maths bounds** (Development) Search plus automatic verification yields new, checkable bounds. [page](https://wheresthe.ai/d/deepmind-alphaevolve-2025-05/) · [source](https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/)
3. **3 Nov 2025: Review of 445 LLM benchmarks finds widespread construct-validity weaknesses** (Development) A reminder that benchmark scores need valid measurement behind them. [page](https://wheresthe.ai/d/oxford-benchmark-construct-validity-2025-11/) · [source](https://arxiv.org/abs/2511.04703)
4. **29 Jan 2026: DeepMind study: most AI 'solutions' to open Erdős problems were already in the literature** (Development) Many 'open' problems fell to literature search, not new ideas. [page](https://wheresthe.ai/d/deepmind-gemini-erdos-case-study-2026-01/) · [source](https://arxiv.org/abs/2601.22401)
5. **20 May 2026: OpenAI model disproves Erdős's 1946 unit-distance conjecture** (Development) First widely accepted historically significant AI result. [page](https://wheresthe.ai/d/openai-unit-distance-disproof-2026-05/) · [source](https://openai.com/index/model-disproves-discrete-geometry-conjecture/)
6. **23 Jul 2026: AI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 2026** (Development) Olympiad maths saturates as a benchmark, including for Chinese labs. [page](https://wheresthe.ai/d/imo-2026-ai-perfect-scores-2026-07/) · [source](https://www.standardmedia.co.ke/sci-tech/article/2001553533/ai-catches-up-with-humans-to-score-100pc-at-top-maths-contest)
7. **4 Sep 2026: Claude completes first end-to-end Lean formal proof of Fermat's Last Theorem** (Development) Machine-checked formalisation at the scale of Wiles's proof. [page](https://wheresthe.ai/d/anthropic-claude-fermat-lean-proof-2026-09/) · [source](https://www.anthropic.com/news/formalizing-fermats-last-theorem)
8. **8 Sep 2026: OpenAI claims Navier-Stokes blow-up proof; mathematicians dispute credit** (Development) The largest claim yet, contested over credit and awaiting acceptance. [page](https://wheresthe.ai/d/openai-navier-stokes-claim-2026-09/) · [source](https://openai.com/index/navier-stokes-solution/)

---
Canonical page: https://wheresthe.ai/line/ai-proves-theorems/ · A Line of Thought from wheresthe.ai
