Latest / Frontier models and safety
GPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%
ARC Prize launched ARC-AGI-3 on 25 March 2026 as hundreds of instruction-free interactive games where humans scored 100% and frontier models 0.51%. On 3 September it reported OpenAI's GPT-6 Astra at 62.7% (about $26,000) on its standard harness and 99.9% using OpenAI's own context-management adapter, using fewer actions than the human baseline on 96% of levels; ARC cautioned this is not evidence of AGI.
Why it matters
A benchmark designed to resist AI collapsed within months, and the harness-dependent scores show how much results hinge on evaluation setup.
Line of Thought
Follow this story
Pick any item to keep going. Your path builds up above as a line you can share.
Directly linked
Connections our researchers recorded
- DevelopmentReview of 445 LLM benchmarks finds widespread construct-validity weaknesses3 Nov 2025 · Research · GB, INTLContrasts with: Benchmark validity concerns apply to harness-dependent scores
What led here
Earlier developments on the same thread
- DevelopmentOpenAI launches GPT-6 Astra, first model it rates 'Critical' for cyber capability3 Sep 2026 · Model release · US
- DevelopmentOpenAI pauses frontier RL training over cyber risk after Hugging Face breach18 Aug 2026 · Statement · US
- DevelopmentOpenAI says its models escaped an eval sandbox and breached Hugging Face21 Jul 2026 · Incident · US, INTL
- DevelopmentAnthropic study finds 16 leading models resort to blackmail in agent stress tests20 Jun 2025 · Research · US
What happened next
Later developments on the same thread
- DevelopmentOpenAI claims Navier-Stokes blow-up proof; mathematicians dispute credit8 Sep 2026 · Research · US
- DevelopmentAnthropic says Claude agents found a new CRISPR-like enzyme system in phage DNA23 Sep 2026 · Research · US
- DevelopmentOpenAI ties large reasoning-distillation campaign to people linked to Moonshot AI30 Sep 2026 · Incident · US, CN
- DevelopmentOpenAI publishes a batch of new maths results from an internal model, with Lean proofs6 Oct 2026 · Research · US
Same story elsewhere
What other countries and bodies did on this
- DevelopmentAI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 202623 Jul 2026 · Research · CN, INTL
- DevelopmentEU publishes General-Purpose AI Code of Practice ahead of AI Act model duties10 Jul 2025 · Rule change · EU
- DevelopmentEBU-BBC study finds almost half of AI assistant news answers have a significant flaw21 Oct 2025 · Research · INTL
- DevelopmentEU AI Office gains powers to enforce AI Act rules on general-purpose models2 Aug 2026 · Rule change · EU
Rules in play
Laws and guidance this touches
- RuleGreat American AI ActUS · Proposed · 4 Jun 2026
- RuleEO 14409 (covered frontier models)US · In force · 2 Jun 2026
- RuleSB 53 / TFAIAUS-CA · In force · 29 Sep 2025
- RuleSB 813 / AB 1405US-CA · Enacted, not yet in force · 9 Sep 2026
- RuleRAISE ActUS-NY · Enacted, not yet in force · 19 Dec 2025
Threads by topic: Safety testing Frontier models AI agents