AI ScoreboardFrontend & Coding Index← Back to rankings

MODEL PROFILE / EDITION 2026.09

Gemini 3.8 Flash

Cost-sensitive teams that need agentic software engineering plus a free experimentation tier

GoogleRank 15Confidence BNo prior snapshot
76.9

Balanced model evidence score; rank 15. No rank delta is shown because there is no prior numeric snapshot.

Best for

Cost-sensitive teams that need agentic software engineering plus a free experimentation tier

Price & access

Standard $0.75 in / $3.75 out per 1M through Dec 31, 2026, then $1.50/$7.50 from Jan 1, 2027; Batch and Flex $0.375/$1.875; Priority $1.35/$6.75; caching $0.075 plus $0.50 per 1M tokens/hour storage

Evidence summary

Pricing, GA date and lifecycle are first-party Google. Context window comes only from Artificial Analysis, modalities are unstated, and the coding numbers come from OpenAI's comparison table rather than Google.

Edition status

No prior snapshot. Tracked in the September 2026 baseline.

Strengths

  • DeepSWE v1.1 73.8% is second-highest in OpenAI's own cross-vendor table, above Claude Fable 5.1 (67.4%)
  • $0.75/$3.75 promotional pricing plus a real free tier
  • Design Arena Website 1301 is the strongest of any sub-$1 input-price model except its own siblings
  • Cyber sibling shows unusually strong vulnerability-patching evidence

Tradeoffs

  • Terminal-Bench 4.0 19.1% in OpenAI's table is far behind frontier agentic scores
  • Arena frontend Elo (1569) sits below the older Gemini 3.7 Flash (1592)
  • Google does not publish context window, max output or modalities on the model or pricing pages
  • AA has no output-speed measurement yet

Score profile

visual 72 · code 73 · agentic 74 · debug 74 · context 86 · speed 80 · value 92

Linked evidence

Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗Evidence source 8 ↗
Compare this model on the full scoreboard

Use the same published data, filters, and side-by-side metrics.

Open comparison view →