Latest / Frontier models and safety

GPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%

ReportFrontierScienceUSConfirmed

ARC Prize launched ARC-AGI-3 on 25 March 2026 as hundreds of instruction-free interactive games where humans scored 100% and frontier models 0.51%. On 3 September it reported OpenAI's GPT-6 Astra at 62.7% (about $26,000) on its standard harness and 99.9% using OpenAI's own context-management adapter, using fewer actions than the human baseline on 96% of levels; ARC cautioned this is not evidence of AGI.

Why it matters

A benchmark designed to resist AI collapsed within months, and the harness-dependent scores show how much results hinge on evaluation setup.

SourceARC PrizeCoverage: ARC Prize (ARC-AGI-3 launch) Checked against the primary source. Independently fact-checked on 7 Oct 2026.
ARC Prize FoundationOpenAI

Line of Thought

Follow this story

Pick any item to keep going. Your path builds up above as a line you can share.

Directly linked

Connections our researchers recorded

What led here

Earlier developments on the same thread

What happened next

Later developments on the same thread

Same story elsewhere

What other countries and bodies did on this

Rules in play

Laws and guidance this touches

Threads by topic: Safety testing Frontier models AI agents