VerdictPal · editorial desk · updated 1 Sep 2026VerdictPal
Glossary

The words the benches use.

Contamination, saturation, pass@1, Elo, ledger — the jargon on benchmark pages, in student English. Named tests live here too, then link out to the full dossier.
Not a leaderboard, not a word salad
A score is a receipt from one test, not a crown. If a word on a benchmark page is still opaque, it should have a card here.

How this site talks

Dossier

tool page · model page

One finite page for one tool or model: dated pricing, privacy, failure modes, and a score you can trace. We never call these cards in reader copy.

Link to this term with the heading.

Status labels

Flagship · Solid · Lightweight · Draft

How finished a public page is. Flagship passed the full gate. Solid is reliable with caveats. Lightweight is a first pass. Draft is scaffolding, not a recommendation.

Link to this term with the heading.

Evidence level

self-run · hands-on · synthesized · vendor claim

How close the desk got to the claim. Self-run means we ran the task. Hands-on means an editor used the tool. Synthesized compiles independent reports. Vendor claim is unverified marketing.

Link to this term with the heading.

Desk

editorial desk · Redaktion

The small Swiss student team that writes, checks dates, and signs off public pages. If the desk has not used a tool, the dossier stays labeled that way.

Link to this term with the heading.

The Pack

newsletter · drop

The editorial newsletter. A drop ships when a tool, model, or method changes — never on a promised weekly cadence.

Link to this term with the heading.

Failure mode

pitfall · what breaks

A documented way the tool or workflow breaks in real student or research use. These stay visible on purpose, even when a vendor would rather hide them.

Link to this term with the heading.

Caveat

blind spot · what it does not measure

The honest limit of a claim. On benchmark pages this is the list of things a high score does not prove.

Link to this term with the heading.

Frontier model

frontier

A current top-tier model, not last year's GPT-4-class leftover. Model compare on this site only lists frontier options.

Link to this term with the heading.

Grounding

source-grounded · cited answer

Tying a claim to a source the reader can open. A fluent paragraph with no trail is not grounded, even if it sounds academic.

Link to this term with the heading.

Hallucination

made-up source · fabricated citation

A confident claim or citation that is not in the source. Fluency is not evidence. Check the PDF before you paste the footnote.

Link to this term with the heading.

Open-weight

open weights · open model

The model weights can be downloaded and run outside the vendor's chat box. Open-weight is not the same as open data, open training, or a privacy guarantee.

Link to this term with the heading.

How scores work

Benchmark

eval · suite · test

A specific test of a specific skill: hard science questions, coding bugs, long documents, agent workflows. It is not a crown for best model.

Link to this term with the heading.

Leaderboard

ranking · scoreboard

A single ordered list of models. VerdictPal refuses that frame. Different tests measure different jobs, so we never collapse Elo, pass@1, and accuracy into one fake ranking.

Link to this term with the heading.

Ledger

score row · score database · evidence ledger

The dated table of sourced scores behind a benchmark page. One row is one model on one test, with a source URL and a date. Unlike metrics stay in their own columns.

Link to this term with the heading.

Source-checked

verified · receipt

VerdictPal opened the cited primary source and confirmed the number. A sourced row without this stamp is still published, but treated with more caution.

Link to this term with the heading.

Self-reported

vendor score · announcement score

The figure comes from the model vendor's own announcement. Useful as a claim, weak as proof. Prefer an independent maintainer or a source-checked row.

Link to this term with the heading.

Snapshot

static snapshot · ingested

A frozen local copy of public scores. The live site does not scrape or call benchmark APIs while you read a page. Dates on the row tell you when that copy was taken.

Link to this term with the heading.

Provenance

source date · ingested at

The paper trail on a score: who published it, the URL, the source date, and when VerdictPal ingested it. No number ships without that trail.

Link to this term with the heading.

Maintainer

official source

The group that owns the test: a paper team, Artificial Analysis, Vals, LMArena, or VerdictPal Lab. Official-maintainer rows beat a third-party re-run of the same suite.

Link to this term with the heading.

Model key

alias · model slug

The stable id VerdictPal uses for one model across sources. Vendors rename display strings; the key and its aliases keep GPT-5.5 from appearing as three different models.

Link to this term with the heading.

Version / variant

edition · ARC-AGI 2

A maintained edition of the same benchmark, such as ARC-AGI-1 versus ARC-AGI-2. Scores from different editions are not interchangeable.

Link to this term with the heading.

pass@1

pass at 1 · pass@k

The share of tasks the model solves on its first try. pass@k would allow k attempts. Coding benches use this because one clean patch matters more than ten lucky retries.

Link to this term with the heading.

Elo

arena elo · rating

A relative rating from pairwise human votes, used by Chatbot Arena. 1200 versus 1300 is a preference gap, not 100 percent accuracy. Never line Elo up against a percent score.

Link to this term with the heading.

Accuracy / percent

pct · % · score

Share of items answered correctly, usually 0–100. High accuracy on a saturated test can mean the suite is too easy, not that the model is done with the job.

Link to this term with the heading.

Human ceiling

expert ceiling · perfect score

The score a careful human or expert panel reaches on the same test. When models meet that line, the suite stops separating the frontier.

Link to this term with the heading.

Headroom

gap to ceiling

The points still left between the best sourced model and the human ceiling. Small headroom means the test is nearly used up.

Link to this term with the heading.

Composite index

intelligence index · Vals Index · AA Index

A weighted blend of several tests from one maintainer. Handy as a glance, opaque as a claim. Always open the underlying suites before you quote the index.

Link to this term with the heading.

Closed-book

no tools · no search

The model must answer from weights alone: no web, no calculator, no files. A closed-book score does not describe a student who is allowed to look things up.

Link to this term with the heading.

Tool-augmented

open-book · with tools · agent harness

The model may use a browser, code runner, or other tools during the test. Compare these rows only to other tool-augmented rows.

Link to this term with the heading.

Harness

agent harness · eval harness

The wrapper that runs the model: prompts, tools, retries, time limits. The same model can look brilliant or clumsy under two harnesses. Read the harness before you trust a coding or agent score.

Link to this term with the heading.

Vendor claim

A number or capability statement that comes from the company selling the product. Published, labeled, and treated as unverified until the desk or an independent source checks it.

Link to this term with the heading.

Artificial Analysis

AA · AA Index

An independent evaluator that re-runs many public suites and publishes its own intelligence index. VerdictPal stores a dated snapshot; we do not call AA while you load the page.

Link to this term with the heading.

Vals AI

Vals Index

Another independent score host. When Vals and the official maintainer both publish the same academic suite, the official row wins on VerdictPal.

Link to this term with the heading.

LMArena

Chatbot Arena · LMSYS

The public human-preference arena behind Elo scores. People vote on anonymous side-by-side answers. It measures taste, not citation quality.

Link to this term with the heading.

Token

The chunk a model reads and writes — roughly a short word or part of a word. Prices are usually per million tokens. Context windows are counted in tokens, not pages.

Link to this term with the heading.

Context window

context length

How much text the model can hold in one prompt, measured in tokens. A huge window does not guarantee the model still uses the sentence on page 40.

Link to this term with the heading.

Retrieval

RAG · search step

Finding the right passage before answering. Long-context tests put the passage in the prompt; search tools must fetch it first. Those are different skills.

Link to this term with the heading.

How tests age

Saturated

used up

Top models cluster near the ceiling. The suite is still useful as history and method, but a new high score barely tells you who is ahead.

Link to this term with the heading.

Solved / defunct

defunct · ~100%

Verified top scores sit at or near 100 percent. We keep the explainer so you can decode old papers, but we do not treat it as live evidence.

Link to this term with the heading.

Deprecated

retired

The maintainer stopped the suite, or the public page is gone. Kept only so old citations still resolve.

Link to this term with the heading.

Contamination

data leak · train/test leak · memorization

Risk that the test items leaked into training data. A high score on a contaminated suite can mean the model saw the answers, not that it learned the skill.

Link to this term with the heading.

Primary benchmark

start here · discriminator

The one live test we would open first in a family. Math's primary is FrontierMath; older MATH and AIME pages stay as solved glossary, not as current evidence.

Link to this term with the heading.

Task families

Family

task family

The job a benchmark claims to measure: knowledge, reasoning, math, coding, agents, and so on. Families keep unlike scores from being stacked into one chart.

Link to this term with the heading.

Knowledge

MMLU

Closed-book recall across school and professional subjects. A high knowledge score is not proof the model can run a literature review.

Link to this term with the heading.

Reasoning

HLE · GPQA

Hard, multi-step questions that should not be solvable by lookup alone. Still an exam, not a measure of messy research productivity.

Link to this term with the heading.

Math

FrontierMath · AIME

Competition or research-level mathematics. Older public suites are largely solved; FrontierMath is the live discriminator on this site.

Link to this term with the heading.

Coding

SWE-bench · HumanEval

Writing or repairing programs. Read the harness: a pass@1 on a hidden unit test is not the same as shipping a reviewed pull request.

Link to this term with the heading.

Agentic

agent · Terminal-Bench

Multi-step work where the model must use tools, a terminal, or a workflow instead of answering one prompt. The harness is half the result.

Link to this term with the heading.

Multimodal

MMMU · vision

Tasks that mix text with images or other media. A strong text-only model can still fail a figure, chart, or slide.

Link to this term with the heading.

Long context

AA-LCR · needle in a haystack

Using evidence that sits far apart inside one very long prompt. This is not the same as searching the live web.

Link to this term with the heading.

Tool use

function calling · BFCL

Whether the model calls the right function, API, or browser action at the right time. Separate from writing good prose about the tool.

Link to this term with the heading.

Safety

HELM Safety

Whether the model refuses harmful requests or leaks things it should not. A safety score is not a privacy policy and not a research-quality grade.

Link to this term with the heading.

Human preference

arena · vibe check

Blind human votes on which answer feels better. Helpful for tone and usefulness. Silent on citations, contamination, and exam skill.

Link to this term with the heading.

Named tests

Every public test on VerdictPal, in one alphabet. Open the dossier for scores, caveats, and sources.

AA output speed

Performance

Output tokens per second · Median output speed

Median output tokens per second after the first chunk arrives, measured on each model's default API provider with a ~1k-token input prompt (AA medium workload).

Abstraction and Reasoning Corpus

Reasoning

ARC Prize · ARC-AGI

Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.

Artificial Analysis Intelligence Index

Reasoning

AA Intelligence Index · AA Index v4.1 · AA Index v4.1.1

AA Intelligence Index v4.1 — weighted composite emphasizing agentic workloads: GDPval-AA v2 (20%), Terminal-Bench 2.1 (16%), τ³-Bench Banking (14%), Humanity's Last Exam (12%), AA-Omniscience (12%), SciCode (8%), GPQA Diamond (6%), AA-LCR (6%), CritPt (6%). IFBench removed for saturation.

CritPt

Reasoning

Critical Physics Tasks · Complex Research using Integrated Thinking - Physics Test

Unpublished research-level physics: 71 composite challenges (70 test + 1 example) written by 50+ active physicists across 11 subfields, plus 190 simpler checkpoint tasks. 6% weight in AA Intelligence Index v4.1.

DeepSWE

Coding

Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.

FrontierMath

Math

FrontierMath Tiers 1-3 · FrontierMath v2

Hard unpublished mathematics: 295 private Tiers 1-3 problems (v2) where the model writes a Python answer() after iterative reasoning and code execution.

GDPval

Agentic

Real-work deliverables vs experts · GDPval-AA v2

GDPval-AA v2 — real knowledge-work deliverables across 44 occupations, graded by blind pairwise comparison against human experts (Elo anchored at 1000). 20% weight in AA Intelligence Index v4.1.

GPQA Diamond

Reasoning

Graduate-Level Google-Proof Q&A

Hard graduate-level science questions (biology, physics, chemistry) written so that non-experts cannot solve them even with web access — the 'Diamond' subset is the highest-quality, expert-validated slice.

HELM Safety

Safety

Holistic Evaluation of Language Models · Stanford HELM

A multi-metric safety evaluation suite from Stanford CRFM that reports model behavior across safety-focused scenarios instead of collapsing everything into one headline score.

HumanEval

Coding

Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.

Humanity's Last Exam

Reasoning

HLE

Cross-domain expert-level questions — math, sciences, humanities, professional law and medicine — designed to be the hardest public exam a frontier model can take.

IFBench

Reasoning

Instruction-Following Benchmark

Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.

LiveBench

Reasoning

A contamination-limited benchmark that refreshes questions over time across reasoning, coding, math, data analysis, language, and instruction-following tasks.

LiveCodeBench

Coding

Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.

LMArena (Chatbot Arena)

Human preference

Chatbot Arena · LMSYS Arena

Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.

LongBench v2

Long context

LongBench 2

Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.

MATH

Math

Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.

MCP Atlas

Tool use

How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.

MMLU-Pro

Knowledge

A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.

MMMU-Pro

Multimodal

MMMU Pro · Massive Multi-discipline Multimodal Understanding Pro

Harder multimodal college exam: original MMMU items with text-only shortcuts filtered out, ten answer options instead of four, and a vision-only setting where the question is inside the image. 3,460 questions across 30 subjects.

Omniscience Accuracy

Knowledge

AA-Omniscience Index · AA-Omniscience

Factual recall and knowledge calibration across 6,000 questions in 42 economically relevant topics and six domains. Accuracy is 8% of AA Intelligence Index v4.1; the non-hallucination slice is another 4%.

SciCode

Coding

Scientific programming tasks requiring domain libraries and multi-step numerical code — AA snapshot runs.

SWE-bench Verified

Coding

Whether a model can resolve real GitHub issues from popular Python repos — it must produce a patch that makes the project's hidden test suite pass. 'Verified' is a 500-issue, human-validated subset of the original SWE-bench.

Terminal-Bench

Agentic

Terminal-Bench 2.1 · [email protected]

Whether an agent can complete realistic, containerized terminal tasks (Terminal-Bench 2.1): debugging code, handling files, running commands, and producing the artifact the hidden checker expects. 16% weight in AA Intelligence Index v4.1.

Terminal-Bench-Science

Agentic

TB-Science · Terminal-Bench-Science 0.1

Whether an agent can finish real scientific research workflows in a terminal: 70 expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, graded on concrete artifacts (analyses, simulations, proofs, code, data products).

Vals Index

Agentic

Vals AI Index

A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.

Vending-Bench 2

Agentic

Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).

τ³-Bench Banking

Agentic

tau3-bench-banking · τ³-Banking

Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.