- Reasoning
- 93
- Coding
- 70
- Agentic
- 100
- Long context
- 70
83average across 4 tested families
Vals Index rows appear first. Models without a Vals row fall back to Artificial Analysis. Source badges show whether a row is single-source or corroborated — not a cross-benchmark winner.
Blended cost per million tokens (input + 3× output) on the X axis, log scale. Intelligence index on the Y axis — Vals when present, otherwise Artificial Analysis. Dots are sized by context window and connected by the Pareto frontier (best intelligence at each cost).
The earlier cut, using input cost alone. Useful when your workload is input-heavy (RAG, document Q&A) rather than output-heavy.
Hover or tab any dot for name, score, and input price.
Older releases, special configurations (Codex CLI, ForgeCode), and ad-hoc aggregations. We surface their sourced scores but we don't promote them to the model atlas until they earn a full model card.
The highest intelligence receipt in each family slice. Use this as a sanity check that a single number isn't hiding a weak spot.
83average across 4 tested families
78average across 4 tested families
91average across 3 tested families
A model with no row in a family is a model that hasn't been publicly tested there. We surface that instead of pretending it's weak.
Every tool card on this list is wired to at least one model in the atlas above. The same wrapper that hides the harness can also hide the price.