Models / OpenAI

GPT-4o (2024-05)

GPT-4o · Released 2024-05-13

OpenAI's May 2024 GPT-4o flagship. MMLU 88.7% and HELM safety 89.4% on public leaderboards. The reference GPT-4o checkpoint before GPT-5.

Legacy OpenAI multimodal baseline — use GPT-5.4/5.5 for new work.

Composite89
Context128K
Input / 1M$2.5
Output / 1M$10
Knowledge cutoff2023-10-01
Statuslightweight
Not enough families covered yet.Add more benchmark rows for GPT-4o (2024-05) to see the radar.
Benchmark placements

Where GPT-4o (2024-05) places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
HELM SafetySafety#4/ 989Stanford HELM leaderboard2024-08-15
Family context

The three highest-scoring carded models in each capability family. Where GPT-4o (2024-05) shows up, it's highlighted.

Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Safety89
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for GPT-4o (2024-05). Scores link to the original source; gaps mean no public row exists yet.

1 sourced rows

1 of 32 catalog benchmarks have a sourced row for GPT-4o (2024-05).

Safety

Sources