MODEL PROFILE / EDITION 2026.09
GLM-5.3
Agentic coding teams that want near-frontier terminal/SWE performance with open weights and low cost
Balanced model evidence score; rank 7. No rank delta is shown because there is no prior numeric snapshot.
Best for
Agentic coding teams that want near-frontier terminal/SWE performance with open weights and low cost
Price & access
$1.40 in / $4.40 out per 1M with $0.26 cached input (Z.ai docs; cached-input storage free for a limited time). The GLM Coding Plan uses points, with 50% off-peak discount outside 14:00-18:00 UTC+8 on weekdays.
Evidence summary
First-party Z.ai launch post and Hugging Face card give detailed benchmarks with context/output settings, and Z.ai's own pricing page gives exact rates. Weakened by the missing SWE-bench-family and Design Arena coverage and by unclear licence text.
Edition status
No prior snapshot. Tracked in the September 2026 baseline.
Strengths
- Terminal Bench 2.1 88.2 is effectively tied with Kimi K3 (88.3) and DeepSeek V4 Pro (87.9)
- Huge generational jump: Terminal Bench 3.0 4.6 -> 28.3 and DeepSWE 46.2 -> 66.9 vs GLM-5.2
- Open weights (753B) with wide inference-engine support and $1.40/$4.40 hosted pricing
- TTFT 2.14 s with 84.2 tok/s - responsive for a frontier-scale model
Tradeoffs
- No SWE-bench Pro, SWE-bench Verified or Design Arena row exists for 5.3
- Thinking cannot be disabled, forcing a migration for non-thinking apps
- Licence terms are not stated on the Hugging Face card; AA calls it a restricted commercial licence
- Text-only on the evidence available - no documented image input for frontend screenshots
Score profile
visual 76 · code 86 · agentic 85 · debug 82 · context 78 · speed 79 · value 88
Linked evidence
Evidence source 1 ↗Evidence source 2 ↗Evidence source 3 ↗Evidence source 4 ↗Evidence source 5 ↗Evidence source 6 ↗Use the same published data, filters, and side-by-side metrics.