Models / StepFun

Step 3.7 Flash

Step · Released 2026-06-01

StepFun's high-efficiency multimodal Flash model, evaluated by Artificial Analysis on 1 Jun. Public catalogs describe image/video input and a 256K context window.

Efficiency-first multimodal model for routing, visual extraction, and cheap agent steps before escalating to a larger model.

#126 of 168 on AA Index · current snapshot
31 AA Index · 10 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite45
Context256K
Input / 1M
Output / 1M
Knowledge cutoffopen
Output speed95 t/s
TTFT (default API)1.72s
Statuslightweight

Family profile

Best published score in each covered benchmark family.

58/100 avg
Reasoning
81
Coding
40
Agentic
39
Long context
70

4 tested benchmark families

Benchmark placements

Where Step 3.7 Flash places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
AA time to first tokenPerformance#35/ 186Artificial Analysis2026-09-01
AA output speedPerformance#43/ 187Artificial Analysis2026-09-01
IFBenchReasoning#63/ 12767Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#84/ 10612Artificial Analysis2026-09-01
AA-LCRLong context#93/ 17970Artificial Analysis2026-09-01
Terminal-BenchAgentic#117/ 18439Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#118/ 18021Artificial Analysis2026-09-01
SciCodeCoding#125/ 18040Artificial Analysis2026-09-01
GPQA DiamondReasoning#136/ 18581Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#145/ 18731Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Step 3.7 Flash shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Step 3.7 Flash in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning81
Coding40
Agentic39
Long context70
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Step 3.7 Flash. Scores link to the original source; gaps mean no public row exists yet.

10 sourced rows

10 of 37 catalog benchmarks have a sourced row for Step 3.7 Flash.

Reasoning

Coding

Agentic

Long context

Performance

Sources