September 6, 2026 · Yotta Labs
GPT-6 Astra Scores 99.9% on ARC-AGI-3 in OpenAI's Own Setup, but 62.7% Under ARC Prize's Independent Environment
My take: The gap between 99.9% and 62.7% on the same benchmark, depending on who runs it, is exactly the kind of detail that matters before accepting any number from a press release.
GPT-6 Astra launched on September 3 with record-setting results across nearly every category: 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4, and 72.6% on OSWorld 2.0, completing computer tasks 47% faster than its predecessor. The issue is that those numbers were generated by OpenAI under its own provider adapter harness. When ARC Prize, the independent organization behind the benchmark, evaluated the same model in its standardized, provider-neutral environment, the result was 62.7%.
That gap does not invalidate the model. But it confirms something worth keeping in mind: when the company that builds the product also runs the benchmark, there is no way to be fully certain the test is free of bias. The 99.9% numbers are worth reading with your own judgment and waiting for independent validation before treating them as definitive.
For those of us using AI models at work, the practical lesson stays the same: benchmarks are a reference point, not a guarantee. How much weight are you giving manufacturer scores versus your own real-world tests?
Want to use these tools? See the unbiased reviews or back to the news.