Latest / Science and research

Review of 445 LLM benchmarks finds widespread construct-validity weaknesses

ResearchScienceFrontierGB INTLConfirmed

A 42-author team with lead authors at the University of Oxford, using 29 expert reviewers, systematically reviewed 445 benchmarks from leading NLP and ML venues and found recurring problems in how phenomena are defined, tasks chosen and scores computed that undermine the claims drawn from them. It sets out eight recommendations for benchmark design and was accepted to the NeurIPS 2025 Datasets and Benchmarks track.

Why it matters

Headline benchmark gains, including in science and maths, are only as meaningful as the measurement behind them, which this review finds often weak.

SourcearXiv (Bean et al., Measuring what Matters) Checked against the primary source. Independently fact-checked on 7 Oct 2026.
University of Oxford

Line of Thought

Follow this story

Pick any item to keep going. Your path builds up above as a line you can share.

Directly linked

Connections our researchers recorded

What led here

Earlier developments on the same thread

What happened next

Later developments on the same thread

Same story elsewhere

What other countries and bodies did on this

Threads by topic: Safety testing Standards and codes