Updated
Updated · arcprize.org · Sep 3
GPT-6 Astra Hits 99.9% on ARC-AGI-3 as It Beats Humans on 96% of Levels
Updated
Updated · arcprize.org · Sep 3

GPT-6 Astra Hits 99.9% on ARC-AGI-3 as It Beats Humans on 96% of Levels

2 articles · Updated · arcprize.org · Sep 3

Summary

  • OpenAI said GPT-6 Astra reached 99.9% on ARC-AGI-3 Semi-Private with its Provider Adapter harness, versus 62.7% under the benchmark’s standard harness.
  • The higher score came with preserved reasoning state and context compaction, which OpenAI said made runs about 3.66 times faster and cut total tokens by 49% across shared solved tasks.
  • On action efficiency, Astra used fewer moves than the median human on 96% of levels and averaged 51.7% fewer actions, a benchmark OpenAI called beyond human parity.
  • Replays showed Astra building compact symbolic world models, inventing domain-specific shorthand to track rules, state and plans; in a sandboxed setup, it also wrote game-specific tools and small software libraries.
  • ARC-AGI-3 is designed to test exploration, modeling, goal-setting and planning in novel environments; OpenAI said saturating the benchmark is not proof of AGI, but a major step toward broader generalization.

Insights

Since GPT-6 Astra already outperforms human efficiency in unfamiliar environments, is the elusive residual gap to true AGI finally closing?
If OpenAI’s GPT-6 Astra can build its own tools to conquer deterministic worlds, what happens when it faces unpredictable real-world chaos?
How did a simple change in memory harness catapult an AI's benchmark score from 62.7% to a near-perfect 99.9% while costing less?

GPT-6 Astra and the 98.6% ARC-AGI-3 Benchmark: Hype, Verification, and the Future of Agentic AI

Overview

OpenAI’s September 2026 launch of GPT-6 Astra marks a major leap in AI, with the model trained on over 100,000 GPUs and leveraging previous AI generations for deeper understanding. Astra’s standout feature is its ability to use computers like a human, automating complex workflows and outperforming rivals on benchmarks. However, its self-reported scores, especially on the ARC-AGI-3 benchmark, remain unverified due to proprietary restrictions, fueling skepticism. Astra’s unprecedented cybersecurity abilities led to a cautious, phased rollout, with advanced features limited to trusted users. This shift is disrupting job markets and forcing new safety and governance strategies as AI enters the AGI era.

...