Atlas
The public catalog of tools and models. It is a desk you can browse, not a ranking of winners.
Dossier · Editorial fit · Open the tool atlas
Link to this term with the heading.
The public catalog of tools and models. It is a desk you can browse, not a ranking of winners.
Dossier · Editorial fit · Open the tool atlas
Link to this term with the heading.
tool page · model page
One finite page for one tool or model: dated pricing, privacy, failure modes, and a score you can trace. We never call these cards in reader copy.
Editorial fit · Failure mode · Evidence level
Link to this term with the heading.
0–100 · fit score · headline score
The 0–100 desk score on a tool dossier. It is a weighted judgment of research-workflow fit, not a user-star average and not a benchmark result.
Status labels · Evidence level · How the 0–100 lands
Link to this term with the heading.
Flagship · Solid · Lightweight · Draft
How finished a public page is. Flagship passed the full gate. Solid is reliable with caveats. Lightweight is a first pass. Draft is scaffolding, not a recommendation.
Editorial fit · Evidence level · Quality gate
Link to this term with the heading.
self-run · hands-on · synthesized · vendor claim
How close the desk got to the claim. Self-run means we ran the task. Hands-on means an editor used the tool. Synthesized compiles independent reports. Vendor claim is unverified marketing.
VerdictPal Lab · Source-checked · How we test
Link to this term with the heading.
editorial desk · Redaktion
The small Swiss student team that writes, checks dates, and signs off public pages. If the desk has not used a tool, the dossier stays labeled that way.
VerdictPal Lab · Evidence level · Who writes this
Link to this term with the heading.
newsletter · drop
The editorial newsletter. A drop ships when a tool, model, or method changes — never on a promised weekly cadence.
Link to this term with the heading.
pitfall · what breaks
A documented way the tool or workflow breaks in real student or research use. These stay visible on purpose, even when a vendor would rather hide them.
Link to this term with the heading.
blind spot · what it does not measure
The honest limit of a claim. On benchmark pages this is the list of things a high score does not prove.
Link to this term with the heading.
first-party · The Lab
Benchmarks we run ourselves, with a frozen protocol, question bank, and published method. Lab scores stay unpublished until a human grades and signs them off.
Citation fidelity · Evidence level · Open the Lab
Link to this term with the heading.
frontier
A current top-tier model, not last year's GPT-4-class leftover. Model compare on this site only lists frontier options.
Link to this term with the heading.
source-grounded · cited answer
Tying a claim to a source the reader can open. A fluent paragraph with no trail is not grounded, even if it sounds academic.
Citation fidelity · Hallucination
Link to this term with the heading.
cited claims
Whether a cited sentence actually appears in the named source. VerdictPal's first-party Lab protocol grades this failure mode for students writing papers.
Grounding · VerdictPal Lab · Hallucination · Citation Fidelity protocol
Link to this term with the heading.
made-up source · fabricated citation
A confident claim or citation that is not in the source. Fluency is not evidence. Check the PDF before you paste the footnote.
Link to this term with the heading.
open weights · open model
The model weights can be downloaded and run outside the vendor's chat box. Open-weight is not the same as open data, open training, or a privacy guarantee.
Link to this term with the heading.
eval · suite · test
A specific test of a specific skill: hard science questions, coding bugs, long documents, agent workflows. It is not a crown for best model.
Leaderboard · Family · Ledger · Benchmark library
Link to this term with the heading.
ranking · scoreboard
A single ordered list of models. VerdictPal refuses that frame. Different tests measure different jobs, so we never collapse Elo, pass@1, and accuracy into one fake ranking.
Link to this term with the heading.
score row · score database · evidence ledger
The dated table of sourced scores behind a benchmark page. One row is one model on one test, with a source URL and a date. Unlike metrics stay in their own columns.
Source-checked · Snapshot · Self-reported
Link to this term with the heading.
verified · receipt
VerdictPal opened the cited primary source and confirmed the number. A sourced row without this stamp is still published, but treated with more caution.
Link to this term with the heading.
vendor score · announcement score
The figure comes from the model vendor's own announcement. Useful as a claim, weak as proof. Prefer an independent maintainer or a source-checked row.
Link to this term with the heading.
static snapshot · ingested
A frozen local copy of public scores. The live site does not scrape or call benchmark APIs while you read a page. Dates on the row tell you when that copy was taken.
Link to this term with the heading.
source date · ingested at
The paper trail on a score: who published it, the URL, the source date, and when VerdictPal ingested it. No number ships without that trail.
Link to this term with the heading.
official source
The group that owns the test: a paper team, Artificial Analysis, Vals, LMArena, or VerdictPal Lab. Official-maintainer rows beat a third-party re-run of the same suite.
Artificial Analysis · Vals AI · LMArena
Link to this term with the heading.
alias · model slug
The stable id VerdictPal uses for one model across sources. Vendors rename display strings; the key and its aliases keep GPT-5.5 from appearing as three different models.
Link to this term with the heading.
edition · ARC-AGI 2
A maintained edition of the same benchmark, such as ARC-AGI-1 versus ARC-AGI-2. Scores from different editions are not interchangeable.
Link to this term with the heading.
pass at 1 · pass@k
The share of tasks the model solves on its first try. pass@k would allow k attempts. Coding benches use this because one clean patch matters more than ten lucky retries.
Link to this term with the heading.
arena elo · rating
A relative rating from pairwise human votes, used by Chatbot Arena. 1200 versus 1300 is a preference gap, not 100 percent accuracy. Never line Elo up against a percent score.
Human preference · Leaderboard
Link to this term with the heading.
pct · % · score
Share of items answered correctly, usually 0–100. High accuracy on a saturated test can mean the suite is too easy, not that the model is done with the job.
Link to this term with the heading.
tps · output speed
How fast a model emits text once it starts answering. Useful for interactive work. It says nothing about whether the answer is right.
Time to first token · Token · Performance
Link to this term with the heading.
TTFT · latency
How long you wait before the first word appears. Separate from tokens-per-second: a model can start slowly and then stream quickly, or the reverse.
Tokens per second · Token · Performance
Link to this term with the heading.
expert ceiling · perfect score
The score a careful human or expert panel reaches on the same test. When models meet that line, the suite stops separating the frontier.
Link to this term with the heading.
gap to ceiling
The points still left between the best sourced model and the human ceiling. Small headroom means the test is nearly used up.
Link to this term with the heading.
intelligence index · Vals Index · AA Index
A weighted blend of several tests from one maintainer. Handy as a glance, opaque as a claim. Always open the underlying suites before you quote the index.
Artificial Analysis · Vals AI · Leaderboard
Link to this term with the heading.
no tools · no search
The model must answer from weights alone: no web, no calculator, no files. A closed-book score does not describe a student who is allowed to look things up.
Link to this term with the heading.
open-book · with tools · agent harness
The model may use a browser, code runner, or other tools during the test. Compare these rows only to other tool-augmented rows.
Closed-book · Agentic · Harness
Link to this term with the heading.
held-out set · hidden test
Some exams keep a hidden question set so models cannot memorize the public list. A score on only the public slice is easier to game.
Contamination · Accuracy / percent
Link to this term with the heading.
agent harness · eval harness
The wrapper that runs the model: prompts, tools, retries, time limits. The same model can look brilliant or clumsy under two harnesses. Read the harness before you trust a coding or agent score.
Agentic · pass@1 · Tool-augmented
Link to this term with the heading.
A number or capability statement that comes from the company selling the product. Published, labeled, and treated as unverified until the desk or an independent source checks it.
Self-reported · Evidence level
Link to this term with the heading.
AA · AA Index
An independent evaluator that re-runs many public suites and publishes its own intelligence index. VerdictPal stores a dated snapshot; we do not call AA while you load the page.
Composite index · Snapshot · Vals AI
Link to this term with the heading.
Vals Index
Another independent score host. When Vals and the official maintainer both publish the same academic suite, the official row wins on VerdictPal.
Artificial Analysis · Maintainer
Link to this term with the heading.
Chatbot Arena · LMSYS
The public human-preference arena behind Elo scores. People vote on anonymous side-by-side answers. It measures taste, not citation quality.
Link to this term with the heading.
The chunk a model reads and writes — roughly a short word or part of a word. Prices are usually per million tokens. Context windows are counted in tokens, not pages.
Context window · Tokens per second
Link to this term with the heading.
context length
How much text the model can hold in one prompt, measured in tokens. A huge window does not guarantee the model still uses the sentence on page 40.
Link to this term with the heading.
RAG · search step
Finding the right passage before answering. Long-context tests put the passage in the prompt; search tools must fetch it first. Those are different skills.
Link to this term with the heading.
The test still separates today's frontier models. This is the set we would open first when you need live evidence.
Saturated · Solved / defunct · Primary benchmark
Link to this term with the heading.
used up
Top models cluster near the ceiling. The suite is still useful as history and method, but a new high score barely tells you who is ahead.
Active · Solved / defunct · Headroom
Link to this term with the heading.
defunct · ~100%
Verified top scores sit at or near 100 percent. We keep the explainer so you can decode old papers, but we do not treat it as live evidence.
Saturated · Deprecated · Human ceiling
Link to this term with the heading.
retired
The maintainer stopped the suite, or the public page is gone. Kept only so old citations still resolve.
Link to this term with the heading.
data leak · train/test leak · memorization
Risk that the test items leaked into training data. A high score on a contaminated suite can mean the model saw the answers, not that it learned the skill.
Public / private split · Saturated
Link to this term with the heading.
start here · discriminator
The one live test we would open first in a family. Math's primary is FrontierMath; older MATH and AIME pages stay as solved glossary, not as current evidence.
Family · Active · Start-here strip
Link to this term with the heading.
task family
The job a benchmark claims to measure: knowledge, reasoning, math, coding, agents, and so on. Families keep unlike scores from being stacked into one chart.
Link to this term with the heading.
MMLU
Closed-book recall across school and professional subjects. A high knowledge score is not proof the model can run a literature review.
Link to this term with the heading.
HLE · GPQA
Hard, multi-step questions that should not be solvable by lookup alone. Still an exam, not a measure of messy research productivity.
Link to this term with the heading.
FrontierMath · AIME
Competition or research-level mathematics. Older public suites are largely solved; FrontierMath is the live discriminator on this site.
Primary benchmark · Solved / defunct
Link to this term with the heading.
SWE-bench · HumanEval
Writing or repairing programs. Read the harness: a pass@1 on a hidden unit test is not the same as shipping a reviewed pull request.
Link to this term with the heading.
agent · Terminal-Bench
Multi-step work where the model must use tools, a terminal, or a workflow instead of answering one prompt. The harness is half the result.
Harness · Tool use · Tool-augmented
Link to this term with the heading.
MMMU · vision
Tasks that mix text with images or other media. A strong text-only model can still fail a figure, chart, or slide.
Link to this term with the heading.
AA-LCR · needle in a haystack
Using evidence that sits far apart inside one very long prompt. This is not the same as searching the live web.
Link to this term with the heading.
function calling · BFCL
Whether the model calls the right function, API, or browser action at the right time. Separate from writing good prose about the tool.
Link to this term with the heading.
speed · latency
Speed and responsiveness: output tokens per second, time to first token. These rows are not quality scores.
Tokens per second · Time to first token
Link to this term with the heading.
HELM Safety
Whether the model refuses harmful requests or leaks things it should not. A safety score is not a privacy policy and not a research-quality grade.
Link to this term with the heading.
arena · vibe check
Blind human votes on which answer feels better. Helpful for tone and usefulness. Silent on citations, contamination, and exam skill.
Link to this term with the heading.
Every public test on VerdictPal, in one alphabet. Open the dossier for scores, caveats, and sources.
Output tokens per second · Median output speed
Median output tokens per second after the first chunk arrives, measured on each model's default API provider with a ~1k-token input prompt (AA medium workload).
TTFT · Median TTFT
Median seconds until the first token returns on the model's default API provider with a ~1k-token input prompt.
ARC Prize · ARC-AGI
Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.
AIME
Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.
AA Intelligence Index · AA Index v4.1 · AA Index v4.1.1
AA Intelligence Index v4.1 — weighted composite emphasizing agentic workloads: GDPval-AA v2 (20%), Terminal-Bench 2.1 (16%), τ³-Bench Banking (14%), Humanity's Last Exam (12%), AA-Omniscience (12%), SciCode (8%), GPQA Diamond (6%), AA-LCR (6%), CritPt (6%). IFBench removed for saturation.
AA-LCR
Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.
Berkeley Tool Calling Leaderboard · BFCL
Whether a model can choose and format function calls correctly across single-turn, multi-turn, parallel, and live API-style tool-use tasks.
Critical Physics Tasks · Complex Research using Integrated Thinking - Physics Test
Unpublished research-level physics: 71 composite challenges (70 test + 1 example) written by 50+ active physicists across 11 subfields, plus 190 simpler checkpoint tasks. 6% weight in AA Intelligence Index v4.1.
Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.
FrontierMath Tiers 1-3 · FrontierMath v2
Hard unpublished mathematics: 295 private Tiers 1-3 problems (v2) where the model writes a Python answer() after iterative reasoning and code execution.
Real-work deliverables vs experts · GDPval-AA v2
GDPval-AA v2 — real knowledge-work deliverables across 44 occupations, graded by blind pairwise comparison against human experts (Elo anchored at 1000). 20% weight in AA Intelligence Index v4.1.
Graduate-Level Google-Proof Q&A
Hard graduate-level science questions (biology, physics, chemistry) written so that non-experts cannot solve them even with web access — the 'Diamond' subset is the highest-quality, expert-validated slice.
Holistic Evaluation of Language Models · Stanford HELM
A multi-metric safety evaluation suite from Stanford CRFM that reports model behavior across safety-focused scenarios instead of collapsing everything into one headline score.
Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.
HLE
Cross-domain expert-level questions — math, sciences, humanities, professional law and medicine — designed to be the hardest public exam a frontier model can take.
Instruction-Following Benchmark
Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.
A contamination-limited benchmark that refreshes questions over time across reasoning, coding, math, data analysis, language, and instruction-following tasks.
Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.
Chatbot Arena · LMSYS Arena
Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.
LongBench 2
Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.
MMMU
College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.
MMLU
Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.
Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.
How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.
A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.
MMMU Pro · Massive Multi-discipline Multimodal Understanding Pro
Harder multimodal college exam: original MMMU items with text-only shortcuts filtered out, ten answer options instead of four, and a vision-only setting where the question is inside the image. 3,460 questions across 30 subjects.
AA-Omniscience Index · AA-Omniscience
Factual recall and knowledge calibration across 6,000 questions in 42 economically relevant topics and six domains. Accuracy is 8% of AA Intelligence Index v4.1; the non-hallucination slice is another 4%.
Scientific programming tasks requiring domain libraries and multi-step numerical code — AA snapshot runs.
Whether a model can resolve real GitHub issues from popular Python repos — it must produce a patch that makes the project's hidden test suite pass. 'Verified' is a 500-issue, human-validated subset of the original SWE-bench.
Terminal-Bench 2.1 · [email protected]
Whether an agent can complete realistic, containerized terminal tasks (Terminal-Bench 2.1): debugging code, handling files, running commands, and producing the artifact the hidden checker expects. 16% weight in AA Intelligence Index v4.1.
TB-Science · Terminal-Bench-Science 0.1
Whether an agent can finish real scientific research workflows in a terminal: 70 expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, graded on concrete artifacts (analyses, simulations, proofs, code, data products).
Vals AI Index
A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.
Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).
tau3-bench-banking · τ³-Banking
Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.