VerdictPal · editorial desk · updated 5 Sep 2026VerdictPal
Benchmarks

Every benchmark, honestly read.

Pick a model and see the public benchmark evidence VerdictPal has found for it. Every row stays dated, sourced, and caveated so the page reads like evidence, not a scoreboard. For the per-model view (ranks, family coverage, source ledger), open the model-first atlas.
This is not a leaderboard
A benchmark is a specific test — coding bugs, hard science questions, agent workflows — not a crown for 'best model.' Scores move, vendors cherry-pick, and old suites leak into training data. Every figure here is dated, sourced, and flagged for contamination risk.

How to read a score

  1. What it measures — the concrete skill or task family this test grades.
  2. What it does not prove — every suite has blind spots; read the caveats before you ship a claim.
  3. When to use it — pick the primary test for your job, then open the dossier for the full ledger.

New to Elo, contamination, or saturation? Open the glossary

27
Public benchmarks
2112
Ledger rows
1965/2112
Scores source-checked
14
Populated sources

The atlas is grouped by what the benchmark claims to measure. Score rows sit underneath each test so unlike metrics never collapse into one fake ranking.

These are the discriminators we would open first in each family. Solved and saturated suites stay in the library below as glossary — not as live evidence.

Omniscience AccuracyAA-Omniscience
KnowledgePrimary · KnowledgeMedium contamination

Use this when · For source-heavy students the failure mode that matters is confident wrong answers. AA-Omniscience is the public suite that jointly scores recall and hallucination across domains that actually show up in papers.

Won't tell you · Open-web research skill — questions are closed-book. A model that should search and cite is not being tested here.

12 score rows · 1 source · refreshed 2026-09-01

Humanity's Last ExamHLE
ReasoningPrimary · ReasoningLow contamination

Use this when · When you need a hard, cross-domain expert exam — math through professional law and medicine — as a closed-book stress test, not a research-productivity proxy.

Won't tell you · Real research productivity — exam-style questions do not capture the messy search-and-synthesize work a researcher actually does.

180 score rows · 1 source · refreshed 2026-09-01

Papers / model cardsFrontierMathFrontierMath Tiers 1-3
MathPrimary · MathLow contamination

Use this when · When MATH and AIME no longer separate the frontier, this is the live math discriminator: original expert problems, Python-enabled, still far from ceiling.

Won't tell you · Tier 4 research-level problems — that 43-problem expansion is a separate, harder set. Do not mix Tiers 1-3 headlines with Tier 4.

15 score rows · 1 source · refreshed 2026-08-15

DeepSWEDeepSWE
CodingPrimary · CodingLow contamination

Use this when · When you need long-horizon, multi-language software work with behavioral verifiers — harder contamination posture than classic SWE-bench Verified.

Won't tell you · Python-only issue fixing like SWE-bench Verified — DeepSWE spans Go, Rust, TypeScript, and more.

8 score rows · 1 source · refreshed 2026-08-04

Terminal-BenchTerminal-Bench 2.1
AgenticPrimary · AgenticLow contamination

Use this when · When you need realistic containerized terminal work — debug, files, commands, artifacts — as an agentic ops signal weighted in AA Index v4.1.

Won't tell you · The model alone — harness, tools, retries, time limits, and container setup can move scores dramatically.

184 score rows · 3 sources · refreshed 2026-09-01

MMMU-ProMMMU Pro
MultimodalPrimary · MultimodalMedium contamination

Use this when · Original MMMU no longer separates the frontier. MMMU-Pro is the live multimodal exam: shortcuts removed, ten options, still a spread below the old 88% ceiling.

Won't tell you · Original MMMU headlines — those still include text-solvable items. A 88%+ MMMU row is not an MMMU-Pro row.

8 score rows · 1 source · refreshed 2026-09-01

Artificial Analysis Long Context ReasoningAA-LCR
Long contextPrimary · Long contextLow contamination

Use this when · Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.

Won't tell you · Retrieval from external databases — the context is in-prompt only.

179 score rows · 1 source · refreshed 2026-09-01

Berkeley BFCLBerkeley Function Calling LeaderboardBFCL
Tool usePrimary · Tool useMedium contamination

Use this when · Tool calling is the plumbing behind agents, automations, and research assistants. BFCL gives that plumbing a public test instead of treating every function-call demo as evidence.

Won't tell you · Whether the downstream tool result is useful — BFCL focuses on choosing and formatting calls, not on the whole product workflow after the call returns.

20 score rows · 1 source · refreshed 2026-06-09

AA output speedOutput tokens per second
PerformancePrimary · PerformanceUnknown contamination

Use this when · When you care how fast tokens stream after the first chunk — IDE feel, chat responsiveness, batch throughput — not whether the answer is right.

Won't tell you · IDE or chat-app latency — only first-party API routes AA tracks as the default provider.

187 score rows · 1 source · refreshed 2026-09-05

Stanford HELMHELM SafetyHolistic Evaluation of Language Models
SafetyPrimary · SafetyMedium contamination

Use this when · Safety evidence is usually marketing copy or red-team anecdotes. HELM gives readers a public, repeatable place to inspect multiple safety metrics side by side.

Won't tell you · Real deployment risk by itself — a model can look safer in HELM than in a product with tools, memory, browsing, or weak policy enforcement.

9 score rows · 1 source · refreshed 2026-06-09

LMArenaLMArena (Chatbot Arena)Chatbot Arena
Human preferencePrimary · Human preferenceLow contamination

Use this when · When you need a human preference signal at scale — how replies feel in blind side-by-side chat — not a correctness exam.

Won't tell you · Correctness — voters reward answers that look good, so formatting, length and confidence can beat accuracy.

13 score rows · 1 source · refreshed 2026-06-09

ActiveSaturatedSolved
Archive · 7 entriesSaturated, solved, or retired
IFBenchInstruction-Following Benchmark
ReasoningSaturatedMedium contamination

Use this when · Historical instruction-following stress test on novel format constraints. Saturated and dropped from AA Intelligence Index v4.1 — glossary honesty, not current frontier ranking.

Won't tell you · Frontier discrimination today — Artificial Analysis removed IFBench from Intelligence Index v4.1 because frontier models saturated it.

127 score rows · 1 source · refreshed 2026-09-01

Papers / model cardsMassive Multitask Language UnderstandingMMLU
KnowledgeSaturatedHigh contamination

Use this when · Historical broad knowledge exam across 57 subjects — useful glossary context for how models were once compared. Saturated at the frontier; not for current ranking.

Won't tell you · Reasoning under uncertainty — a model can pattern-match the right letter without understanding the question.

22 score rows · 2 sources · refreshed 2026-06-09

Papers / model cardsMassive Multi-discipline Multimodal UnderstandingMMMU
MultimodalSaturatedMedium contamination

Use this when · When you need college-level multimodal understanding — charts, diagrams, structures, scans — not text-only knowledge exams.

Won't tell you · Pure vision quality — many items can be partly solved from the text alone, inflating scores.

8 score rows · 1 source · refreshed 2026-06-09

SWE-benchSWE-bench Verified
CodingSaturatedMedium contamination

Use this when · Once the flagship real-issue Python coding test — still useful lineage. Frontier scores now cluster tightly; prefer DeepSWE or SWE-Lancer for current agent signal.

Won't tell you · Anything but Python — the suite is Python-only, so it says little about other languages or front-end work.

23 score rows · 2 sources · refreshed 2026-07-01

American Invitational Mathematics ExaminationAIME
MathDefunctMedium contamination

Use this when · Historical contest-math signal — once useful for hard integer reasoning. Solved for frontier models with extended thinking; keep for context, not current ranking.

Won't tell you · Tiny sample (15 problems/year) — one problem is ~6.7%, so single-run scores are noisy; demand pass@k or multiple seeds.

37 score rows · 3 sources · refreshed 2026-09-01

Vendor claimsHumanEval
CodingDefunctHigh contamination

Use this when · Historical code-generation baseline — short Python functions from docstrings. Solved and heavily contaminated; keep for lineage, not frontier coding rank.

Won't tell you · Real software work — 164 self-contained toy functions are nothing like a production codebase.

15 score rows · 3 sources · refreshed 2026-06-09

MATH
MathDefunctHigh contamination

Use this when · Historical competition-math suite — once a standard reasoning yardstick. Now saturated; keep for context, prefer harder math evals for frontier separation.

Won't tell you · Whether the reasoning is correct — only the final answer is graded, so a right answer from wrong work still scores.

25 score rows · 3 sources · refreshed 2026-09-01

Evidence console — model lookup (14 sources)

Look up model evidence

Pick one or two models from the ledger. Scores stay grouped by benchmark so an Arena Elo never gets ranked against MMLU accuracy.

Last ledger refresh: 2026-09-05Static snapshot; sources linked belowComparing 2 / 2 model slots
Model evidence brief — 61 rows
61 rows found
OpenAI

GPT 5.5

Released 2026-04-23

31 rows
Knowledge
Omniscience Accuracy

AA-Omniscience Accuracy

57%
Artificial Analysis — AA-OmniscienceSource-checkedRank 5

Source· source 2026-09-01· ingested 2026-09-01

MMLU-Pro

MMLU-Pro

81.6%
MMLU-Pro HuggingFace leaderboardNeeds auditRank 2

Source· source 2026-04-23· ingested 2026-06-09

Reasoning
CritPt

CritPt composite (AA run)

27.1%
Artificial Analysis — CritPtSource-checkedRank 7

Source· source 2026-09-01· ingested 2026-09-01

GPQA Diamond

GPQA Diamond accuracy (CoT)

81.2%
GPQA Diamond paper leaderboardSource-checkedRank 3

Source· source 2026-05-15· ingested 2026-06-05

IFBench

IFBench (AA run)

75.85%
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

LiveBench

LiveBench

69.8%
LiveBench leaderboardNeeds auditRank 3

Source· source 2026-04-15· ingested 2026-06-09

Math
FrontierMath

FrontierMath Tiers 1-3 (v2) best Epoch run

51.7%
Epoch AI FrontierMath Tiers 1-3 (v2) CSVSource-checkedRank 2

Source· source 2026-08-15· ingested 2026-08-15

MATH

MATH

89.6%
OpenAI model release notesNeeds auditRank 2

Source· source 2026-04-23· ingested 2026-06-09

Coding
DeepSWE

DeepSWE pass@1 (xhigh effort)

67%
DeepSWE leaderboard v1.1Source-checkedRank 2

Source· source 2026-06-20· ingested 2026-06-22

HumanEval

HumanEval

96.2% pass@1
CodeSOTA HumanEval leaderboardNeeds auditRank 2

Source· source 2026-04-23· ingested 2026-06-09

SciCode

SciCode (AA run)

56.1%
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

SWE-bench Verified

SWE-bench Verified

82.5%
SWE-bench Verified leaderboardSource-checked

Source· source 2026-06-01· ingested 2026-06-09

Agentic
GDPval

GDPval-AA v2

1531
Artificial Analysis — GDPval-AA v2Source-checkedRank 3

Source· source 2026-06-17· ingested 2026-06-22

τ³-Bench Banking

τ³-Bench Banking (AA run)

38.97%
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

Terminal-Bench

Terminal-Bench 2.0 (audited harness)

82%
Terminal-Bench 2.0 leaderboardSource-checkedRank 2

Source· source 2026-05-20· ingested 2026-06-05

Vals Index

Vals Index

68%
Vals AI — Vals IndexSource-checked

Source· source 2026-06-04· ingested 2026-06-09

Vending-Bench 2

Final bank balance (USD)

7523.84
Andon Labs Vending-Bench 2Source-checkedRank 6

Source· source 2026-08-04· ingested 2026-08-04

Multimodal
Long context
LongBench v2

LongBench v2

56.1%
LongBench v2 leaderboardNeeds auditRank 3

Source· source 2026-04-01· ingested 2026-06-09

Tool use
MCP Atlas

MCP Atlas

77.8%
MCP Atlas leaderboardNeeds auditRank 3

Source· source 2026-06-09· ingested 2026-06-09

Performance
AA output speed

Median output tokens/s (1k prompt, default provider)

0 t/s
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

AA time to first token

Median time to first token (1k prompt, default provider)

0.00s
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

Safety
HELM Safety

HELM Safety

90.6%
Stanford HELM leaderboardNeeds auditRank 3

Source· source 2026-04-23· ingested 2026-06-09

Human preference
Anthropic

Claude Opus 4.8

Released 2026-05-28

30 rows
Knowledge
Omniscience Accuracy

AA-Omniscience Accuracy

48.8%
Artificial Analysis — AA-OmniscienceSource-checkedRank 18

Source· source 2026-09-01· ingested 2026-09-01

MMLU-Pro

MMLU-Pro

82.4%
MMLU-Pro HuggingFace leaderboardNeeds auditRank 1

Source· source 2026-05-28· ingested 2026-06-09

Reasoning
CritPt

CritPt composite (AA run)

20.9%
Artificial Analysis — CritPtSource-checkedRank 11

Source· source 2026-09-01· ingested 2026-09-01

GPQA Diamond

GPQA Diamond accuracy (CoT)

79.8%
GPQA Diamond paper leaderboardSource-checkedRank 4

Source· source 2026-05-15· ingested 2026-06-05

IFBench

IFBench (AA run)

62.24%
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

LiveBench

LiveBench

70.8%
LiveBench leaderboardNeeds auditRank 2

Source· source 2026-04-15· ingested 2026-06-09

Math
FrontierMath

FrontierMath Tiers 1-3 (v2) best Epoch run

47.24%
Epoch AI FrontierMath Tiers 1-3 (v2) CSVSource-checkedRank 5

Source· source 2026-08-15· ingested 2026-08-15

MATH

MATH

88.4%
Anthropic Claude 4 familyNeeds auditRank 3

Source· source 2026-05-28· ingested 2026-06-09

Coding
DeepSWE

DeepSWE pass@1 (max effort)

59%
DeepSWE leaderboard v1.1Source-checkedRank 3

Source· source 2026-06-20· ingested 2026-06-22

HumanEval

HumanEval

95.8% pass@1
Anthropic Claude 4 familyNeeds auditRank 3

Source· source 2026-05-28· ingested 2026-06-09

SciCode

SciCode (AA run)

53.5%
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

SWE-bench Verified

SWE-bench Verified

88.6%
SWE-bench official leaderboardSource-checkedRank 2

Source· source 2026-05-30· ingested 2026-06-05

Agentic
GDPval

GDPval-AA v2

1638
Artificial Analysis — GDPval-AA v2Source-checkedRank 2

Source· source 2026-06-17· ingested 2026-06-22

τ³-Bench Banking

τ³-Bench Banking (AA run)

34.23%
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

Terminal-Bench

Terminal-Bench 2.0 (audited harness)

74.6%
Terminal-Bench 2.0 leaderboardSource-checkedRank 5

Source· source 2026-05-20· ingested 2026-06-05

Terminal-Bench-Science

TB-Science 0.1 resolution (Claude Code)

10.5%
Terminal-Bench-Science 0.1 announcementSource-checkedRank 4

Source· source 2026-08-27· ingested 2026-09-01

Vals Index

Vals Index

70.4%
Vals AI — Vals IndexSource-checked

Source· source 2026-06-04· ingested 2026-06-09

Multimodal
Long context
LongBench v2

LongBench v2

58.4%
LongBench v2 leaderboardNeeds auditRank 1

Source· source 2026-04-01· ingested 2026-06-09

Tool use
MCP Atlas

MCP Atlas

79.2%
MCP Atlas leaderboardNeeds auditRank 2

Source· source 2026-06-09· ingested 2026-06-09

Performance
AA output speed

Median output tokens/s (1k prompt, default provider)

0 t/s
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

AA time to first token

Median time to first token (1k prompt, default provider)

0.00s
Artificial AnalysisSource-checked

Source· source 2026-09-01· ingested 2026-09-01

Safety
HELM Safety

HELM Safety

91.8%
Stanford HELM leaderboardNeeds auditRank 2

Source· source 2026-05-28· ingested 2026-06-09

Human preference
Score database — 2112 ledger rows

This is the sorting database behind the brief. It includes the reviewed snapshot plus legacy score rows from benchmark pages.

1807

api · api

Vals AI23

manual-snapshot · manual

SWE-bench22

official-json · scheduled

Terminal-Bench25

official-page · manual

LMArena13

official-page · manual

CodeSOTA6

manual-snapshot · manual

Stanford HELM9

official-page · manual

Berkeley BFCL20

official-page · manual

LongBench19

official-page · manual

LiveBench21

official-page · manual

Papers / model cards95

manual-snapshot · manual

Vendor claims34

manual-snapshot · manual

DeepSWE8

official-page · manual

Vending-Bench10

official-page · manual

ModelRelease dateBenchmarkScoreWithin-benchmark scaleSourceDate
A.X-K2Other2026-08-12AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Agnes 2.5 Pro AlphaOther2026-07-24AA output speedPerformance · Active174 t/sArtificial Analysis2026-09-01
Agnes 2.5 Pro BetaOther2026-08-26AA output speedPerformance · Active152 t/sArtificial Analysis2026-09-01
Apodex 1.1Other2026-08-30AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude 3 OpusAnthropic2024-03-04AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude 4 Opus (Reasoning)Anthropic2025-05-22AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude 4.1 Opus (Reasoning)Anthropic2025-08-05AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-29AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Fable 5Anthropic2026-06-09AA output speedPerformance · Active58 t/sArtificial Analysis2026-09-01
Claude Opus 4.5Anthropic2025-11-24AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 4.5 (Reasoning)Anthropic2025-11-24AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 4.6Anthropic2026-02-05AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-05AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 4.7Anthropic2026-04-16AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-16AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 4.8Anthropic2026-05-28AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic2026-07-24AA output speedPerformance · Active47 t/sArtificial Analysis2026-09-01
Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic2026-07-24AA output speedPerformance · Active47 t/sArtificial Analysis2026-09-01
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic2026-07-24AA output speedPerformance · Active49 t/sArtificial Analysis2026-09-01
Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic2026-07-24AA output speedPerformance · Active46 t/sArtificial Analysis2026-09-01
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic2026-07-24AA output speedPerformance · Active48 t/sArtificial Analysis2026-09-01
Claude Sonnet 4.6Anthropic2026-02-17AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-17AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic2026-02-17AA output speedPerformance · Active45 t/sArtificial Analysis2026-09-01
Claude Sonnet 5Anthropic2026-06-30AA output speedPerformance · Active80 t/sArtificial Analysis2026-09-01
Command ACohere2026-03-15AA output speedPerformance · Active228 t/sArtificial Analysis2026-09-01
DeepSeek Coder V2DeepSeek2024-05-15AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
DeepSeek R1DeepSeek2025-01-20AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
DeepSeek V3.2 (Reasoning)DeepSeek2025-12-01AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
DeepSeek V4 Flash (Non-reasoning)DeepSeek2026-04-24AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek2026-04-24AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek2026-07-31AA output speedPerformance · Active106 t/sArtificial Analysis2026-09-01
DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek2026-08-21AA output speedPerformance · Active111 t/sArtificial Analysis2026-09-01
DeepSeek V4 Pro (Non-reasoning)DeepSeek2026-04-24AA output speedPerformance · Active50 t/sArtificial Analysis2026-09-01
DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek2026-04-24AA output speedPerformance · Active50 t/sArtificial Analysis2026-09-01
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek2026-08-13AA output speedPerformance · Active51 t/sArtificial Analysis2026-09-01
DiffusionGemma 26B A4BGoogle2026-06-10AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
EXAONE 4.5 33BOther2026-04-09AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
G9v3-39A5BOther2026-08-03AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
G9v3-3BOther2026-07-23AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 1.0 UltraGoogle2024-02-08AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3 Flash PreviewGoogle2025-12-17AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3 Flash Preview (Reasoning)Google2025-12-17AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3 ProGoogle2025-11-18AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3 Pro Preview (low)Google2025-11-18AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3.1 Flash LiteGoogle2026-05-07AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3.1 Pro PreviewGoogle2026-02-19AA output speedPerformance · Active113 t/sArtificial Analysis2026-09-01
Gemini 3.5 FlashGoogle2026-05-19AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
Gemini 3.5 Flash (medium)Google2026-05-19AA output speedPerformance · Active232 t/sArtificial Analysis2026-09-01
Gemini 3.5 Flash (minimal)Google2026-05-19AA output speedPerformance · Active217 t/sArtificial Analysis2026-09-01
Gemini 3.5 Flash-LiteGoogle2026-07-21AA output speedPerformance · Active316 t/sArtificial Analysis2026-09-01
Gemini 3.6 Flash (high)Google2026-07-21AA output speedPerformance · Active161 t/sArtificial Analysis2026-09-01
Gemini 3.7 FlashGoogle2026-08-13AA output speedPerformance · Active307 t/sArtificial Analysis2026-09-01
Gemma 4 12B (Non-reasoning)Google2026-06-10AA output speedPerformance · Active130 t/sArtificial Analysis2026-09-01
Gemma 4 12B (Reasoning)Google2026-06-07AA output speedPerformance · Active131 t/sArtificial Analysis2026-09-01
Gemma 4 31B ITGoogle2026-03-01AA output speedPerformance · Active36 t/sArtificial Analysis2026-09-01
GLM 5V Turbo (Reasoning)Zhipu2026-04-01AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
GLM-4.7 (Reasoning)Zhipu2025-12-22AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
GLM-5 (Non-reasoning)Zhipu2026-02-11AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
GLM-5 (Reasoning)Zhipu2026-02-11AA output speedPerformance · Active0 t/sArtificial Analysis2026-09-01
The Lab

Public scores show what vendors choose to publish. The Lab is where VerdictPal runs its own tests, with method, prompts, and raw examples so you can repeat them.

Open the Lab