85.6%
MATH
Models / Anthropic
Claude Opus · Released 2024-07-01
July 2024 Claude Opus 4 checkpoint. HELM safety 92.1% on the public ledger. Predecessor to Opus 4.5+.
Early Opus 4 API snapshot for leaderboard continuity.
Where Claude Opus 4 (2024-07) places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
| Benchmark | Family | Rank | Score | Source | Date |
|---|---|---|---|---|---|
| HELM Safety | Safety | #1/ 9 | 92 | Stanford HELM leaderboard | 2024-08-15 |
| MATH | Math | #14/ 25 | 86 | Anthropic model card | 2025-06-01 |
The three highest-scoring models with pages in each capability family. Where Claude Opus 4 (2024-07) shows up, it's highlighted.
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.
Every catalog benchmark for Claude Opus 4 (2024-07). Scores link to the original source; gaps mean no public row exists yet.
2 of 37 catalog benchmarks have a sourced row for Claude Opus 4 (2024-07).
85.6%
MATH
92.1%
HELM Safety