Review of 445 LLM benchmarks finds widespread construct-validity weaknesses
A 42-author team with lead authors at the University of Oxford, using 29 expert reviewers, systematically reviewed 445 benchmarks from leading NLP and ML venues and found recurring problems in how phenomena are defined, tasks chosen and scores computed that undermine the claims drawn from them. It sets out eight recommendations for benchmark design and was accepted to the NeurIPS 2025 Datasets and Benchmarks track.
Why it matters
Headline benchmark gains, including in science and maths, are only as meaningful as the measurement behind them, which this review finds often weak.
Line of Thought
Follow this story
Pick any item to keep going. Your path builds up above as a line you can share.
Curated lines through this story
Directly linked
Connections our researchers recorded
- DevelopmentGPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%3 Sep 2026 · Report · USContrasts with: Benchmark validity concerns apply to harness-dependent scores
What led here
Earlier developments on the same thread
- DevelopmentGemini Deep Think earns officially graded gold-medal score at IMO 202521 Jul 2025 · Research · GB, US
What happened next
Later developments on the same thread
- DevelopmentDeepMind study: most AI 'solutions' to open Erdős problems were already in the literature29 Jan 2026 · Research · GB, US
- DevelopmentAI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 202623 Jul 2026 · Research · CN, INTL
Same story elsewhere
What other countries and bodies did on this
- DevelopmentSouth Korea's AI Basic Act takes effect, with a one-year pause on fines22 Jan 2026 · Rule change · KR
- DevelopmentEU publishes General-Purpose AI Code of Practice ahead of AI Act model duties10 Jul 2025 · Rule change · EU
Threads by topic: Safety testing Standards and codes