AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

GPT-5.6 Luna

Cost-sensitive teams running lots of code edits and CI-style fixes

OpenAIRank 25Confidence BNo prior snapshot
72.5

Balanced model evidence score; rank 25. No rank delta is shown because there is no prior numeric snapshot.

Best for

Cost-sensitive teams running lots of code edits and CI-style fixes

Price & access

$0.20 in / $1.20 out per 1M; cached input $0.02

Evidence summary

First-party pricing/context plus third-party SWE-bench Pro; frontend and agentic evidence is thin and partly family-level.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • SWE-bench Pro 62.7% at $0.20/$1.20 is an extreme price/quality ratio
  • 1.05M context at budget pricing
  • 127.2 tok/s throughput
  • Cached input $0.02 makes repeat-context agents very cheap

Tradeoffs

  • Arena frontend Elo 1516 is near the bottom of this cohort
  • TTFT 124.70 s at max effort
  • AA Intelligence Index 43 is well below the frontier
  • No visual/design benchmark evidence

Score profile

visual 62 · code 76 · agentic 70 · debug 68 · context 86 · speed 66 · value 93

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →