MODEL PROFILE / EDITION 2026.09
GPT-6 Astra
Teams that want the single strongest one-shot frontend generator and can absorb frontier pricing
Balanced model evidence score; rank 4. No rank delta is shown because there is no prior numeric snapshot.
Best for
Teams that want the single strongest one-shot frontend generator and can absorb frontier pricing
Price & access
$10 in / $50 out per 1M; cached input $1.00, cache write $12.50; above 272K context $20 in / $75 out; Fast mode 2x price; Batch $5/$25
Evidence summary
Pricing, context and modality data are first-party (OpenAI docs/pricing/launch post). Frontend rank is a fresh Arena snapshot with only 1,055 votes, and the two AA Intelligence Index figures in circulation (55 vs 61.2) conflict.
Edition status
No prior snapshot. Tracked in the September 2026 baseline.
Strengths
- Top of Arena WebDev Frontend by a clear margin (1808)
- Best published Terminal-Bench 4.0 (57.9%) and long-context retrieval (MRCR v2 96.3% at 512K-1M)
- Strongest screen/UI grounding evidence: ScreenSpot-Pro 92.7%, OSWorld 2.0 72.6%
- Lowest published hallucination rate in its cohort (4.2%)
Tradeoffs
- Most expensive model here ($10/$50, doubling above 272K context)
- Very high latency at max effort (AA TTFT 317 s)
- No SWE-bench Pro/Verified numbers published
- Arena rank rests on only 1,055 votes (+/-26), so rank 1 vs 2 is inside the interval
Score profile
visual 96 · code 92 · agentic 96 · debug 94 · context 86 · speed 42 · value 34
Linked evidence
Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗Use the same published data, filters, and side-by-side metrics.