AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

Grok 4.6

Interactive, visually ambitious app building inside Cursor or Grok Build

xAI (listed as 'SpaceXAI' on Arena, AA and its own docs)Rank 16Confidence BNo prior snapshot
76.5

Balanced model evidence score; rank 16. No rank delta is shown because there is no prior numeric snapshot.

Best for

Interactive, visually ambitious app building inside Cursor or Grok Build

Price & access

$2 in / $6 out per 1M with $0.50 cached input (75% discount); xAI charges different rates for requests exceeding the 200K context window and the exact surcharge is not published; fast variant is 2x price

Evidence summary

Two human-preference boards, first-party docs for context and price, and AA measurement - but zero published SWE-bench-family scores and an undisclosed long-context surcharge.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • Arena rank 11 (1637) and Design Arena rank 10 (1312) - consistently strong on both preference boards
  • $2/$6 with 75%-off cached input is aggressive for a frontier model
  • Explicitly tuned for long-running agents and 'more ambitious interactive and visual work'
  • First-class availability in Cursor, Vercel and Cloudflare

Tradeoffs

  • No published SWE-bench Pro/Verified/LiveCodeBench score at all
  • 500K context is half the field's 1M norm, and pricing above 200K is undisclosed
  • AA index drops from 51 to 42 across effort tiers
  • TTFT 52 s at high effort

Score profile

visual 79 · code 79 · agentic 80 · debug 76 · context 74 · speed 62 · value 78

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →