Why this benchmark is useful
Safety evidence is usually marketing copy or red-team anecdotes. HELM gives readers a public, repeatable place to inspect multiple safety metrics side by side.
Benchmarks / Safety
Holistic Evaluation of Language Models
A multi-metric safety evaluation suite from Stanford CRFM that reports model behavior across safety-focused scenarios instead of collapsing everything into one headline score.
Safety evidence is usually marketing copy or red-team anecdotes. HELM gives readers a public, repeatable place to inspect multiple safety metrics side by side.
Read HELM as a dashboard. Do not average the metrics into a VerdictPal-style winner; look for the specific failure class your workflow cannot tolerate.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4 (2024-07)AnthropicHELM safety scenario accuracy; composite across harmlessness and helpfulness metrics. | 2024-07-01 | 92.1% | Stanford HELM leaderboard2024-08-15 | Needs audit |
| 2 | Claude Opus 4.8AnthropicHELM safety scenario accuracy; composite across harmlessness and helpfulness metrics. | 2026-05-28 | 91.8% | Stanford HELM leaderboard2026-05-28 | Needs audit |
| 3 | GPT 5.5OpenAIHELM safety scenario accuracy. | 2026-04-23 | 90.6% | Stanford HELM leaderboard2026-04-23 | Needs audit |
| 4 | GPT-4o (2024-05)OpenAIHELM safety scenario accuracy. | 2024-05-13 | 89.4% | Stanford HELM leaderboard2024-08-15 | Needs audit |
| 5 | Claude Sonnet 4.6AnthropicHELM safety scenario accuracy. | 2026-02-17 | 89.2% | Stanford HELM leaderboard2026-02-17 | Needs audit |
| 6 | Gemini 3.1 ProGoogleHELM safety scenario accuracy. | 2026-02-19 | 88.9% | Stanford HELM leaderboard2026-02-19 | Needs audit |
| 7 | Gemini 1.5 Pro (2024-06)GoogleHELM safety scenario accuracy. | 2024-06-01 | 88.7% | Stanford HELM leaderboard2024-08-15 | Needs audit |
| 8 | DeepSeek V4 Pro MaxDeepSeekHELM safety scenario accuracy. | 2026-04-24 | 86.4% | Stanford HELM leaderboard2026-04-24 | Needs audit |
| 9 | Mistral Large 2 (2024-07)MistralHELM safety scenario accuracy. | 2024-07-24 | 84.3% | Stanford HELM leaderboard2024-08-15 | Needs audit |
Primary source is the Stanford CRFM HELM site. VerdictPal should add rows only when the exact scenario, metric, model, and HELM release are preserved.
A strong HELM safety slice does not prove safe tool use, safe retrieval, or safe institutional deployment.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.