Models / OpenAI

GPT-4 (2023)

GPT-4 · Released 2023-03-14

Original GPT-4 technical report baseline. MMLU 86.4%, HumanEval 67.0% pass@1. Anchor for pre-2024 frontier comparisons.

Historical anchor only — not deployable on current OpenAI routes.

Composite67
Context8K
Input / 1M$30
Output / 1M$60
Knowledge cutoff2021-09-01
Statuslightweight
KnowledgeMathCoding
Knowledge
86
Math
53
Coding
67

69average across 3 tested families

Benchmark placements

Where GPT-4 (2023) places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
HumanEvalCoding#15/ 1567GPT-4 Technical Report2023-03-15
MMLUKnowledge#17/ 2286GPT-4 Technical Report2023-03-15
MATHMath#20/ 2053GPT-4 Technical Report2023-03-15
Family context

The three highest-scoring carded models in each capability family. Where GPT-4 (2023) shows up, it's highlighted.

Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge86
Math53
Coding67
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for GPT-4 (2023). Scores link to the original source; gaps mean no public row exists yet.

3 sourced rows

3 of 32 catalog benchmarks have a sourced row for GPT-4 (2023).

Knowledge

Math

Coding

Sources