Updated September 26, 2026 · 6 min read
GPT-6 Astra benchmark results, explained
Astra posts eye-watering numbers — 99.9% on ARC-AGI-3, 97.6% on FrontierMath. But benchmarks measure narrow things, and knowing which score maps to which real-world job is the difference between a smart purchase and an expensive mistake.
The full scoreboard
| Benchmark | What it measures | Astra | Comparison |
|---|---|---|---|
| OSWorld 2.0 | Real computer-use tasks | 72.6% | Sol 65.7%, in 47% less time |
| Terminal-Bench 4.0 | Agentic coding in a terminal | 57.9% | Sol 37.3% |
| FrontierMath Tier 4 v2 | Research-level mathematics | 97.6% | Sol 83.0%, Fable 5.1 87.8% |
| ARC-AGI-3 | Abstract reasoning / novel problems | 99.9% | — |
| ExploitBench | Finding & exploiting vulnerabilities | 100% | First model to max it |
| BenchCAD | 3D object reconstruction from renders | 95.9% | Sol 83.3%, Fable 84.3% |
| AA Intelligence Index | Broad general intelligence | 61 | Sol 61, Fable 5.1 66 |
Three numbers that actually matter
1. OSWorld 2.0 — 72.6%
This is the one to watch. It measures whether a model can genuinely operate a computer: navigating apps, completing workflows. Astra beats Sol's score while taking ~47% less time per task. If you care about AI doing work autonomously, this is your metric.
2. ExploitBench — 100%
Astra is the first model to max it out, which is why OpenAI rates it "Critical" in its Preparedness Framework and gates exploit development behind a vetted defensive-security program. Impressive — and a reminder that capability and risk travel together.
3. AA Intelligence Index — 61
The reality check. On broad general intelligence, Astra ties the previous Sol generation and trails Claude Fable 5.1 by five points. The frontier moved forward in specific directions — math, tool use, computer operation — not across the board.
What the benchmarks don't tell you
- Benchmarks aren't your workflow. ExploitBench is irrelevant if you're planning a birthday party; FrontierMath won't help you draft emails.
- Reliability isn't scored. Early users report Astra still fumbles simple tasks occasionally — a 99.9% ARC score doesn't mean 99.9% reliability on your errands.
- Safety trade-offs are real. OpenAI itself flagged that Astra's chain of thought is harder to monitor than earlier models.
How to use benchmarks wisely
The honest method: take 3–5 tasks you actually do, run them on Astra and on your current model, and compare. Thirty minutes of your own testing beats any leaderboard. Our review does exactly this for common professional work.