AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

Grok 4.5

Teams that need a documented SWE-bench Pro number from xAI's line rather than 4.6's unpublished claims

xAI (listed as 'SpaceXAI')Rank 20Confidence BNo prior snapshot
74.8

Balanced model evidence score; rank 20. No rank delta is shown because there is no prior numeric snapshot.

Best for

Teams that need a documented SWE-bench Pro number from xAI's line rather than 4.6's unpublished claims

Price & access

$2 in / $6 out per 1M; cached input $0.30 (BenchLM notes this is a 10% cache-read estimate, not an xAI-published rate); higher-context pricing applies above 200K with an undisclosed surcharge

Evidence summary

A hard third-party SWE-bench Pro number plus two preference boards and first-party pricing, but the release date and AA index both have conflicting published values.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • SWE-bench Pro 64.7% (rank 10) beats GPT-5.6 Sol's 64.6% and Opus 4.7's 64.3%
  • Roughly 2x token efficiency and under half the steps of comparable models per xAI
  • Well-sampled Arena row (6,034 votes) and Design Arena 1301
  • Same $2/$6 price as Grok 4.6

Tradeoffs

  • Superseded by Grok 4.6, which scores higher on both preference boards at the same price
  • AA publishes two conflicting Intelligence Index values (45 vs 56)
  • 500K context with an undisclosed >200K surcharge
  • Cached-input rate is an estimate, not an xAI-published price

Score profile

visual 71 · code 80 · agentic 78 · debug 74 · context 74 · speed 68 · value 78

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →