Models / Other

Muse Spark

Muse · Released 2026-01-01

Muse's tax-domain model. TaxEval v2 leader at 77.7% on Vals. Niche vertical — not a general-purpose frontier model.

Tax and compliance workflows only — do not use as a general research or coding model.

#41 of 168 on AA Index · current snapshot
44 AA Index · 10 benchmark rows · 2 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite60
Context128K
Input / 1M
Output / 1M
Knowledge cutoff2025-06-01
Statuslightweight

Family profile

Best published score in each covered benchmark family.

64/100 avg
Reasoning
88
Math
39
Coding
52
Agentic
62
Long context
77

5 tested benchmark families

Benchmark placements

Where Muse Spark places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
FrontierMathMath#9/ 1539Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
IFBenchReasoning#21/ 12776Artificial Analysis2026-09-01
AA-LCRLong context#29/ 17977Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#32/ 18041Artificial Analysis2026-09-01
SciCodeCoding#35/ 18052Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#52/ 18744Artificial Analysis2026-09-01
GPQA DiamondReasoning#61/ 18588Artificial Analysis2026-09-01
Terminal-BenchAgentic#62/ 18462Artificial Analysis2026-09-01
AA time to first tokenPerformance#176/ 186Artificial Analysis2026-09-01
AA output speedPerformance#177/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Muse Spark shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Muse Spark in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning88
Math39
Coding52
Agentic62
Long context77
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Muse Spark. Scores link to the original source; gaps mean no public row exists yet.

10 sourced rows

10 of 37 catalog benchmarks have a sourced row for Muse Spark.

Reasoning

Math

Coding

Agentic

Long context

Performance

Sources