86.4%
MMLU
Models / OpenAI
GPT-4 · Released 2023-03-14
Original GPT-4 technical report baseline. MMLU 86.4%, HumanEval 67.0% pass@1. Anchor for pre-2024 frontier comparisons.
Historical anchor only — not deployable on current OpenAI routes.
69average across 3 tested families
Where GPT-4 (2023) places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
| Benchmark | Family | Rank | Score | Source | Date |
|---|---|---|---|---|---|
| HumanEval | Coding | #15/ 15 | 67 | GPT-4 Technical Report | 2023-03-15 |
| MMLU | Knowledge | #17/ 22 | 86 | GPT-4 Technical Report | 2023-03-15 |
| MATH | Math | #20/ 20 | 53 | GPT-4 Technical Report | 2023-03-15 |
The three highest-scoring carded models in each capability family. Where GPT-4 (2023) shows up, it's highlighted.
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover or tab any dot for name, score, and input price.
Every catalog benchmark for GPT-4 (2023). Scores link to the original source; gaps mean no public row exists yet.