MODEL PROFILE / EDITION 2026.09
Claude Opus 4.8
Backend-heavy repo work where SWE-bench-style correctness matters more than UI taste
Balanced model evidence score; rank 18. No rank delta is shown because there is no prior numeric snapshot.
Best for
Backend-heavy repo work where SWE-bench-style correctness matters more than UI taste
Price & access
$5 in / $25 out per 1M; cache read $0.50, cache write $6.25
Evidence summary
Bedrock card plus Anthropic pricing, two human-preference boards, and multiple vendor comparison tables that agree on its tier.
Edition status
No prior snapshot. Tracked in the September 2026 baseline.
Strengths
- SWE-bench Verified 88.6% and Pro 69.2% remain top-10 results
- NL2Repo 69.7 is the highest in DeepSeek's and Z.ai's published comparison tables
- Deeply sampled Arena row (9,804 votes)
- Lower TTFT (42 s) than the Fable tier
Tradeoffs
- Design Arena Website 1269 (rank 33) shows weak visual preference relative to its code strength
- Legacy status on Anthropic's pricing page
- Same $5/$25 price as the much stronger Opus 5
- Arena effort aliases differ by 23 Elo
Score profile
visual 68 · code 85 · agentic 83 · debug 80 · context 92 · speed 60 · value 55
Linked evidence
Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗Compare this model on the full scoreboard
Open comparison view →Use the same published data, filters, and side-by-side metrics.