Benchmarks / Safety

HELM Safety

Holistic Evaluation of Language Models

A multi-metric safety evaluation suite from Stanford CRFM that reports model behavior across safety-focused scenarios instead of collapsing everything into one headline score.

SafetySolidActiveMedium contamination riskSince 2022
What this does not measure
  • Real deployment risk by itself — a model can look safer in HELM than in a product with tools, memory, browsing, or weak policy enforcement.
  • Institution-specific policy fit — campus, clinical, legal, and workplace rules often require narrower tests than a public benchmark can encode.
  • A single rank — HELM is deliberately multi-metric, so precision, robustness, toxicity, fairness, calibration, and efficiency should be read separately.
Analysis

Why this benchmark is useful

Safety evidence is usually marketing copy or red-team anecdotes. HELM gives readers a public, repeatable place to inspect multiple safety metrics side by side.

Scope

Coverage map

Task family
Safety
Format
Scenario-based evaluations across safety-relevant tasks and metrics; leaderboard pages expose metric families rather than one universal score.
Scoring
Multiple metrics by scenario. Read the exact metric label and HELM release before comparing models.
Maintainer
Stanford CRFM
Reading guide

How to read the scores

Read HELM as a dashboard. Do not average the metrics into a VerdictPal-style winner; look for the specific failure class your workflow cannot tolerate.

Blind spots

What it does not cover

  • Real deployment risk by itself — a model can look safer in HELM than in a product with tools, memory, browsing, or weak policy enforcement.
  • Institution-specific policy fit — campus, clinical, legal, and workplace rules often require narrower tests than a public benchmark can encode.
  • A single rank — HELM is deliberately multi-metric, so precision, robustness, toxicity, fairness, calibration, and efficiency should be read separately.
Scores

Evidence ledger

9 rows
99 rows
92.1%best score
0source-checked
1sources
2026-05-28to 2024-08-15
Stanford HELM9

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.891.8% · Stanford HELM leaderboard
  2. DeepSeek V4 Pro Max86.4% · Stanford HELM leaderboard
  3. GPT 5.590.6% · Stanford HELM leaderboard
  4. Gemini 3.1 Pro88.9% · Stanford HELM leaderboard
  5. Claude Sonnet 4.689.2% · Stanford HELM leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4 (2024-07)AnthropicHELM safety scenario accuracy; composite across harmlessness and helpfulness metrics.2024-07-0192.1%
Stanford HELM leaderboard2024-08-15
Needs audit
2Claude Opus 4.8AnthropicHELM safety scenario accuracy; composite across harmlessness and helpfulness metrics.2026-05-2891.8%
Stanford HELM leaderboard2026-05-28
Needs audit
3GPT 5.5OpenAIHELM safety scenario accuracy.2026-04-2390.6%
Stanford HELM leaderboard2026-04-23
Needs audit
4GPT-4o (2024-05)OpenAIHELM safety scenario accuracy.2024-05-1389.4%
Stanford HELM leaderboard2024-08-15
Needs audit
5Claude Sonnet 4.6AnthropicHELM safety scenario accuracy.2026-02-1789.2%
Stanford HELM leaderboard2026-02-17
Needs audit
6Gemini 3.1 ProGoogleHELM safety scenario accuracy.2026-02-1988.9%
Stanford HELM leaderboard2026-02-19
Needs audit
7Gemini 1.5 Pro (2024-06)GoogleHELM safety scenario accuracy.2024-06-0188.7%
Stanford HELM leaderboard2024-08-15
Needs audit
8DeepSeek V4 Pro MaxDeepSeekHELM safety scenario accuracy.2026-04-2486.4%
Stanford HELM leaderboard2026-04-24
Needs audit
9Mistral Large 2 (2024-07)MistralHELM safety scenario accuracy.2024-07-2484.3%
Stanford HELM leaderboard2024-08-15
Needs audit
Method

What it covers

Primary source is the Stanford CRFM HELM site. VerdictPal should add rows only when the exact scenario, metric, model, and HELM release are preserved.

Score ceiling

Where it breaks down

A strong HELM safety slice does not prove safe tool use, safe retrieval, or safe institutional deployment.

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.