MODEL PROFILE / EDITION 2026.09
DeepSeek V4 Flash
High-concurrency, low-latency coding assistance and batch page generation on a tight budget
Balanced model evidence score; rank 14. No rank delta is shown because there is no prior numeric snapshot.
Best for
High-concurrency, low-latency coding assistance and batch page generation on a tight budget
Price & access
Peak $0.44 in / $1.32 out per 1M; off-peak exactly half at $0.22/$0.66; cache-hit input $0.014 peak / $0.007 off-peak. Peak hours 01:00-04:00 and 06:00-10:00 UTC Mon-Fri.
Evidence summary
First-party pricing and version table, first-party weights card with harness disclosure, two preference boards, two benchmark leaderboards plus an independent Qwen re-run, and AA speed measurements.
Edition status
No prior snapshot. Tracked in the September 2026 baseline.
Strengths
- Fastest time-to-first-token measured here (0.95 s) with 131.8 tok/s throughput
- Terminal-Bench 2.1 82.7% and SWE-bench Verified 79.0% at $0.44/$1.32
- MIT-licensed 304B weights - practical to self-host on a single high-end node
- 2,500-request concurrency limit, 5x the V4 Pro tier
Tradeoffs
- Design Arena Website 1223-1255 puts visual taste in the bottom third
- SWE-bench Pro 52.6% is mid-pack
- Text-only unless you switch to the vision-exp sibling
- Two different Design Arena rows for near-identical names invite confusion
Score profile
visual 69 · code 75 · agentic 74 · debug 72 · context 82 · speed 86 · value 96
Linked evidence
Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Evidence source 7 ↗Evidence source 8 ↗Evidence source 9 ↗Use the same published data, filters, and side-by-side metrics.