Models / MiniMax

MiniMax M2.5

MiniMax · Released 2026-04-01

MiniMax's April 2026 agentic model. SWE-bench Verified 80.2% (#8). Superseded by M3 for open-weights work but still on leaderboards.

Prior MiniMax tier for coding agents when pinned to M2.5 endpoints.

#94 of 168 on AA Index · current snapshot
35 AA Index · 10 benchmark rows · 2 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite55
Context1M
Input / 1M$0.30
Output / 1M$1.2
Knowledge cutoff2026-02-01
Statussolid

Family profile

Best published score in each covered benchmark family.

68/100 avg
Reasoning
85
Coding
80
Agentic
35
Long context
72

4 tested benchmark families

Benchmark placements

Where MiniMax M2.5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
SWE-bench VerifiedCoding#10/ 2380SWE-bench official leaderboard2026-05-30
IFBenchReasoning#44/ 12772Artificial Analysis2026-09-01
AA-LCRLong context#75/ 17972Artificial Analysis2026-09-01
GPQA DiamondReasoning#98/ 18585Artificial Analysis2026-09-01
SciCodeCoding#99/ 18043Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#118/ 18735Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#119/ 18021Artificial Analysis2026-09-01
Terminal-BenchAgentic#135/ 18435Artificial Analysis2026-09-01
AA time to first tokenPerformance#171/ 186Artificial Analysis2026-09-01
AA output speedPerformance#172/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where MiniMax M2.5 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for MiniMax M2.5 in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning85
Coding80
Agentic35
Long context72
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for MiniMax M2.5. Scores link to the original source; gaps mean no public row exists yet.

10 sourced rows

10 of 37 catalog benchmarks have a sourced row for MiniMax M2.5.

Reasoning

Coding

Agentic

Long context

Performance

Sources