Claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 don’t hold up to scrutiny

1 hour ago 17

A claim circulating on social media that OpenAI’s GPT-6 Astra scored 98.6% on the ARC-AGI-3 benchmark would, if true, represent one of the most significant leaps in AI capability ever recorded. The problem: there’s no verified evidence to back it up.

The ARC-AGI-3 benchmark, which launched on March 25, 2026, is specifically designed to test whether AI models can navigate interactive environments without instructions or predefined objectives. When models first encountered the benchmark, scores came in below 1%.

What the leaderboard actually shows

The current top performer on ARC-AGI-3 is Anthropic’s Claude Opus 5, sitting at roughly 30.2%.

OpenAI’s own GPT-5.6 Sol, its most recent model with verified benchmark results, managed 7.78% under official testing conditions. When OpenAI used its custom Responses API settings, that number climbed to 38.3%.

No confirmed ARC-AGI-3 scores exist for a model called Astra. As of early September 2026, OpenAI hasn’t even established a confirmed release date or official branding for GPT-6.

What we actually know about Astra

OpenAI previewed Astra around August 1, 2026. The model earned recognition for solving approximately 10 open math problems, which are unsolved challenges that have stumped mathematicians.

The gap between Astra’s demonstrated capabilities and the claimed 98.6% score suggests one of two scenarios. Either the claim conflates or inflates GPT-5.6 Sol’s results, or it originates from an unverified source making assertions about Astra’s performance without supporting data.

Why ARC-AGI-3 matters

The ARC Prize Foundation designed ARC-AGI-3 as a successor to earlier versions of the benchmark, each iteration becoming harder to game through memorization or pattern-matching. The benchmark uses a scoring methodology called Relative Human Action Efficiency (RHAE), which allows for a comparative evaluation of AI models against human baselines in interactive settings.

When initial models scored below 1%, it underscored just how difficult this benchmark is. The progression to Claude Opus 5’s 30.2% over roughly six months represents meaningful but incremental progress.

The competitive landscape and market implications

Anthropic’s Claude Opus 5 holds the current ARC-AGI-3 lead, while OpenAI’s GPT-5.6 Sol shows competitive performance under optimized conditions.

Astra’s mathematical reasoning achievements are noteworthy, and OpenAI’s standing in ARC-AGI-3 specifically, based on available data, places it behind Anthropic’s leading model.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article