AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

GPT-6 Astra

Teams that want the single strongest one-shot frontend generator and can absorb frontier pricing

OpenAIRank 4Confidence BNo prior snapshot
82.4

Balanced model evidence score; rank 4. No rank delta is shown because there is no prior numeric snapshot.

Best for

Teams that want the single strongest one-shot frontend generator and can absorb frontier pricing

Price & access

$10 in / $50 out per 1M; cached input $1.00, cache write $12.50; above 272K context $20 in / $75 out; Fast mode 2x price; Batch $5/$25

Evidence summary

Pricing, context and modality data are first-party (OpenAI docs/pricing/launch post). Frontend rank is a fresh Arena snapshot with only 1,055 votes, and the two AA Intelligence Index figures in circulation (55 vs 61.2) conflict.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • Top of Arena WebDev Frontend by a clear margin (1808)
  • Best published Terminal-Bench 4.0 (57.9%) and long-context retrieval (MRCR v2 96.3% at 512K-1M)
  • Strongest screen/UI grounding evidence: ScreenSpot-Pro 92.7%, OSWorld 2.0 72.6%
  • Lowest published hallucination rate in its cohort (4.2%)

Tradeoffs

  • Most expensive model here ($10/$50, doubling above 272K context)
  • Very high latency at max effort (AA TTFT 317 s)
  • No SWE-bench Pro/Verified numbers published
  • Arena rank rests on only 1,055 votes (+/-26), so rank 1 vs 2 is inside the interval

Score profile

visual 96 · code 92 · agentic 96 · debug 94 · context 86 · speed 42 · value 34

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →