AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

Claude Opus 4.8

Backend-heavy repo work where SWE-bench-style correctness matters more than UI taste

AnthropicRank 18Confidence ANo prior snapshot
75.2

Balanced model evidence score; rank 18. No rank delta is shown because there is no prior numeric snapshot.

Best for

Backend-heavy repo work where SWE-bench-style correctness matters more than UI taste

Price & access

$5 in / $25 out per 1M; cache read $0.50, cache write $6.25

Evidence summary

Bedrock card plus Anthropic pricing, two human-preference boards, and multiple vendor comparison tables that agree on its tier.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • SWE-bench Verified 88.6% and Pro 69.2% remain top-10 results
  • NL2Repo 69.7 is the highest in DeepSeek's and Z.ai's published comparison tables
  • Deeply sampled Arena row (9,804 votes)
  • Lower TTFT (42 s) than the Fable tier

Tradeoffs

  • Design Arena Website 1269 (rank 33) shows weak visual preference relative to its code strength
  • Legacy status on Anthropic's pricing page
  • Same $5/$25 price as the much stronger Opus 5
  • Arena effort aliases differ by 23 Elo

Score profile

visual 68 · code 85 · agentic 83 · debug 80 · context 92 · speed 60 · value 55

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →