EDITION 2026.09 / EDITORIAL EVIDENCE INDEX
Two contests.
One scoreboard.
Tools are the products you work in — builders, agentic IDEs, design-to-code. Models are the engines underneath them. They are scored in two separate contests, because a workflow and a foundation model are not the same unit.
Editorial evidence score, not a lab benchmark. No hands-on build was performed. How scoring works →
HOW TO READ THIS BOARD
Twelve seconds. Then everything else on the page makes sense.
3 QUESTIONS / NO ACCOUNT
Find my tool
Answer three plain questions. The finder re-weights the published evidence model and returns starting points with the reason attached — never a single "winner".
Pick an answer to the first question to see recommendations.
TOOLS LANE / PRODUCT EVIDENCE MODEL
The tools index
20 ranked tools shown
| Rank | Tool | Score | Dimension profile | Why it ranks here | Price / free tier | Code export | Best for | Conf. | Compare |
|---|
Tracked, not ranked
These entries are in the index and searchable, but the rubric evidence needed for an ordinal rank does not exist yet. No score is invented, and published ranks are untouched.
Ranked entries are commercial or open products with first-party frontend evidence. Terminal-only agents, spec-driven PR bots, and brand-new releases enter as tracked entries first. Nominate a missing tool by opening an issue against the published method.
INTERACTIVE MODEL
Weighting Lab
Change what matters and the board re-ranks instantly. The model stays constrained to 100%, and the preset is written into the URL so you can share the exact view you are looking at.
Overall = Σ (subscore × dimension weight ÷ 100)
Whole-number subscores; overall rounded to one decimal. Other sliders normalise proportionally when one changes. Published ranks always use the Balanced preset.
CATEGORY LENS
Category champions
Different jobs, different winners. Category context prevents false equivalence.
MODEL RANKING LENSES
Quality, purpose, and economic efficiency
The Quality / Cost lens prices a standard coding task, applies a nonlinear anti-junk gate, and reports value logarithmically so +10 points means 10× better value.
QUALITY / COST FRONTIER
The Frontier Atlas
Up is better. Left is cheaper. The models on the glowing edge deliver quality that cheaper rivals cannot match within the evidence tolerance.
100 UNDERLYING MODELS / 30 EVIDENCE-RANKED
The models index
Loading 100 models…
Tracked model catalogue
These models complete the Top 100 catalogue. They have verified availability metadata, but no ordinal quality rank until benchmark evidence is normalized across comparable harnesses.
CATALOGUE / NOT A COMMON RANKING
Every entry in the index
Tools and models are grouped by type. Each score is shown only inside its own evidence model.
A product workflow and an underlying foundation model are different units. Combining them would create false precision, so the catalogue keeps them apart and labels the type on every card.
Frontend & coding tools product evidence model
Underlying models balanced model evidence
SIDE BY SIDE
Compare
Pick two or three entries in the same lane. The comparison is written into the URL, so you can send it to someone else.
WHAT CHANGED / EDITION 2026.09
Updates
Every edition states what moved and why. Subscribe with the Atom feed.
Market log / 2025—2027
METHOD / SOURCE QUALITY
Facts are linked. Judgments are labelled.
The index uses separate evidence models for products and foundation models. Type is always visible because tools and models are not comparable units.
Tool formula
0.20·VF + 0.20·ED + 0.15·SP + 0.15·IT + 0.10·EC + 0.10·DP + 0.10·VP
Subscores are whole numbers from 0–100. Weighted overall scores are rounded to one decimal.
Model formulas
Quality = 0.25·VI + 0.20·CO + 0.15·AG + 0.10·DE + 0.10·CT + 0.10·SP + 0.10·VA
Quality / Cost prices 1M input tokens at a 70% cache-hit share plus 200K output, applies a confidence-adjusted logistic quality gate, and reports quality-adjusted cost in decibels. +10 points means 10× better value.
Evidence hierarchy
Official product, pricing, and documentation pages first; primary financial disclosures second; credible reporting and independent hands-on evaluations for corroboration and limitations.
Provisional policy
An entry is TRACKED when it is real and relevant but lacks the frontend-specific evidence the rubric requires. Tracked entries get no ordinal rank, no invented subscores, and are excluded from every ranked count.
Missing facts stay n.a.; conflicts stay visible. Corrections are welcome and are logged in the next edition's Updates section.