AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

Claude Opus 5

Production teams that want near-flagship correctness at half of Fable pricing

AnthropicRank 2Confidence ANo prior snapshot
84.0

Balanced model evidence score; rank 2. No rank delta is shown because there is no prior numeric snapshot.

Best for

Production teams that want near-flagship correctness at half of Fable pricing

Price & access

$5 in / $25 out per 1M; cache read $0.50, cache write $6.25

Evidence summary

First-party launch post and Bedrock card, Anthropic pricing, two human-preference boards and two third-party benchmark leaderboards agree.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • SWE-bench Verified 96% is the single highest coding score in this evidence base
  • Best AA Coding Agent Index v1.4 (68.1) of any model in OpenAI's own table
  • Half the token price of Fable 5.1 with a large, well-sampled Arena row (8,431 votes)
  • Documented frontend QA behaviour: opened desktop and phone widths, found and fixed an off-screen checkout button

Tradeoffs

  • $5/$25 is still premium pricing
  • TTFT 77.82 s at max effort
  • Fable 5.1 beats it on SWE-bench Pro, Terminal-Bench 4.0, OSWorld and Design Arena
  • Effort tier changes its Arena rank by 28 Elo (max vs high)

Score profile

visual 88 · code 96 · agentic 92 · debug 88 · context 92 · speed 55 · value 55

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →