AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

Kimi K3

Design-led teams and self-hosters who want the best visual-preference model with open weights

Moonshot AIRank 3Confidence ANo prior snapshot
83.8

Balanced model evidence score; rank 3. No rank delta is shown because there is no prior numeric snapshot.

Best for

Design-led teams and self-hosters who want the best visual-preference model with open weights

Price & access

$3 in / $15 out per 1M with $0.30 cached input (Moonshot); Alibaba resells at $2.827/$14.133 in most regions and $3/$15 in Singapore/US international; OpenRouter median $2.55/$12.75 with $0.256 cache read

Evidence summary

First-party Hugging Face card, two independent human-preference boards (one #1 placement), Alibaba and OpenRouter price rows, and AA measurements. Only the AA index value is disputed.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • Design Arena Website rank 1 (1361) - the best pure design-preference score in this set
  • Arena WebDev Frontend rank 5 (1691) - top open-weight model on that board after Qwen3.8 Max
  • Open weights at 2.8T total / 104B active with a 1M context and native vision
  • MCPMark-Verified 94.5 and Toolathlon 76.5 make it a genuinely strong tool user

Tradeoffs

  • Self-hosting means a ~1.56 TB download and multi-node serving
  • $3/$15 API pricing is high for an open-weight model
  • AA measured only 40.3 tok/s (OpenRouter providers reach 90)
  • AA index figures in circulation conflict (50 at max vs 57 claimed at launch)

Score profile

visual 88 · code 88 · agentic 87 · debug 82 · context 92 · speed 64 · value 74

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗Evidence source 8 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →