Matharena

Matharena

Free benchmark leaderboard that scores LLMs on elite competition math from AIME and IMO to Lean proof datasets

65/100MonitorFreeFree

If your question is "which model actually reasons through hard math," MathArena is the cleanest free answer available, and the clickable raw-output cells let you audit the score instead of trusting it. The September 11 entry for Mistral Prover (Leanstral 1.5 + K3) at 38/46 verified solutions on ArXivLean June 2026 (82.61%) is a good example of the specificity you get. Its narrowness is the feature: no composite index hiding failure modes. Just don't bring it a question about coding or conversation — it won't answer one, and its API-free, read-only design means you can't automate the benchmark run.

Verified 16h ago · liveness 65/100 · cite: rightaichoice.com/tools/matharena

Best for
  • AI researchers measuring mathematical reasoning in frontier LLMs
  • Evaluation engineers auditing raw model outputs
  • Competition-math enthusiasts comparing models
  • ML teams selecting a math-capable model before a STEM deployment
Not ideal for
  • Teams evaluating coding, general knowledge, or conversational ability
  • Anyone needing an API or automated benchmarking pipeline
  • Buyers who need multimodal evaluation beyond static math problems
Visit Website

AdvancedZero setup: MathArena is a public website, so researchers get value in under a minute by picking a dataset month and reading the table. Evaluation engineers auditing outputs add a couple of minutes per cell to review raw model text. No install, no account, no API key.WebNo public APIVerified 16h ago
Pricing
Free
FreeFree tier
Learning curve
Advanced
Zero setup: MathArena is a public website, so researchers get value in under a minute by picking a dataset month and reading the table. Evaluation engineers auditing outputs add a couple of minutes per cell to review raw model text. No install, no account, no API key.
Runs on
Web
No public API
Who it's for
AI researcherEvaluation engineerCompetition-math enthusiast
Live sentiment
Is Matharena actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip MathArena if you need coding, chat, or general-knowledge model scores, an API to run benchmarks in your own pipeline, or any multimodal evaluation beyond the deprecated Kangaroo math set.

The 30-second take
Price reality

MathArena is completely free — $0/mo for full leaderboard access, with no paid tiers, seat limits, or usage caps. That makes it cheaper than commercial eval platforms like Artificial Analysis or LMSYS-adjacent paid dashboards. The trade-off is scope: you get free math-competition rankings only, and no automation or API, so the real cost is your manual reading time.

In short

Matharena — Free benchmark leaderboard that scores LLMs on elite competition math from AIME and IMO to Lean proof datasets. Best for AI researchers measuring mathematical reasoning in frontier LLMs, Evaluation engineers auditing raw model outputs, Competition-math enthusiasts comparing models. Free to use.

What's new in Matharena

Checked 14 days ago

Across the latest 4 updates: 4 feature updates.

What people actually say about Matharena — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

39 mentions across 2 sources (Hacker News, GitHub) · researched Jul 3, 2026.

68% positive32% critical

Average across the 2 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Uses fresh, uncontaminated competition problems for honest evaluation.
  • +Transparent per-cell raw output viewing for detailed analysis.
  • +Covers final-answer, proof-based, and visual math benchmarks.
  • +Frequent updates with latest competition results, staying current.
  • +Free and open platform with community-submitted leaderboards.
Recurring frustrations
  • Reproducibility is inconsistent — some models get wildly different scores.
  • Documentation is sparse, confusing setup for new users.
  • Only GPT-5 (High) gets Agent mode, unfair for open models.
  • Requested models (Gemini Flash 2.5, Claude Opus 4.6) not added promptly.
  • No official API or automation support for CI/CD integration.
Patterns worth knowing
Contamination-free benchmarks are highly valued — fresh competition problems ensure scores reflect real reasoning, not memorization.
Seen on Hacker News
Reproducibility problems plague the evaluation pipeline, with users getting inconsistent results from the same code.
Seen on GitHub
Open-source models are gaining ground on proprietary ones, with StepFun-3.5 topping AIME 2026, exciting the community.
Seen on Hacker News
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • No hidden costs, but you need your own API keys to evaluate models, which can be expensive for large runs.

Viability Score

65/100
Monitor

How well maintained and how widely used is Matharena? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
68
What the vendor publishes
0

Last calculated: September 2026

How we score →

Key Features

  • Interactive LLM math benchmark leaderboard
  • Dataset track: ArXivLean with monthly views (03/2026 through 06/2026)
  • Dataset track: BrokenArXiv with monthly views (02/2026 through 06/2026 and Overall)
  • Dataset track: ArXivMath with monthly views (12/2025 through 06/2026 and Overall)
  • Click any leaderboard cell to view the raw model output
  • Filter leaderboards by competition or dataset month
  • Compare models side-by-side across entries
  • Deprecated: Visual Math via Kangaroo 2025 grades 1-2 through 11-12 and Overall
  • Deprecated: Final-Answer Comps including AIME 2025/2026, HMMT Feb 2025, HMMT Nov 2025, HMMT Feb 2026, BRUMO 2025, SMT 2025, CMIMC 2025, Apex, Apex Shortlist
  • Deprecated: Proof-Based Comps including USAMO 2025/2026, IMO 2025, IMC 2025, Miklós Schweitzer 2025, Putnam 2025
  • Deprecated: Project Euler problem view
  • Mistral Prover (Leanstral 1.5 + K3) scored 38/46 verified on ArXivLean June 2026 (82.61%)
  • Muse Spark 1.3 added September 13
  • Claude-Opus-5 (max) added to leaderboard
  • Kimi K3 and Muse Spark 1.1 added to leaderboard

About Matharena

FreeAdvancedNo APIWeb

MathArena is a free leaderboard that tests large language models on hard competition mathematics, built by SRI Lab at ETH Zurich and INSAIT. Evaluation runs across three dataset tracks — ArXivLean, BrokenArXiv, and ArXivMath — with month-by-month views tracking each release; the site's original visual math (Kangaroo grades 1-12), final-answer contest (AIME, HMMT, BRUMO, SMT, CMIMC, Apex), proof-based (USAMO, IMO, Putnam, IMC, Miklós Schweitzer), and Project Euler views are now labeled deprecated in favor of the dataset tracks. The workflow is filter-first: pick a dataset or competition month, read the table, then click any cell to see the raw model output behind the score — so you can judge whether a win came from real reasoning or a lucky final answer. The board moves quickly: Mistral Prover (Leanstral 1.5 + K3) landed September 11 with 38/46 verified solutions on ArXivLean June 2026 (82.61%), Muse Spark 1.3 followed September 13, and earlier entries include Claude-Opus-5 (max), Kimi K3, and Muse Spark 1.1. It is built for AI researchers, evaluation engineers, and competition-math people who need a narrow, citable read on mathematical reasoning — not coding, chat, or general-knowledge scores.

Behind the Verdict

MathArena's value comes from what it refuses to do. It measures one thing — how well frontier LLMs handle elite competition mathematics — and it measures it three ways: ArXivLean, BrokenArXiv, and ArXivMath dataset tracks with month-level granularity, plus a set of legacy views (Visual Math Kangaroo, Final-Answer Comps, Proof-Based Comps, Project Euler) the site now labels deprecated. The September 11 addition of Mistral Prover (Leanstral 1.5 + K3), with 38/46 verified solutions on ArXivLean June 2026 (82.61%), and the September 13 addition of Muse Spark 1.3 show the cadence — new models arrive within days of release. Earlier entries like Claude-Opus-5 (max), Kimi K3, and Muse Spark 1.1 round out the comparison set. The strongest design choice is the raw-output cell. A leaderboard gives you a rank; MathArena lets you click in and read the actual argument the model produced. That's the difference between a score you trust and a score you can check. For evaluation engineers auditing proof-based work, this is the closest thing to a free instrument that exists. Where it falls short is scope and automation. It is mathematical competition problems only — no coding, no chat, no general knowledge, no multimodal evaluation beyond the deprecated Kangaroo set. There is no mention of an API, so you cannot wire MathArena into a benchmarking pipeline; every run is a manual, browser-based read. Server slowdowns are possible since it's a free academic resource. Use it as a reference and a sanity check, not as your CI benchmark harness.

Researching Matharena? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Matharena actually fits — and what changes day-one when you adopt it.

AI researcher

You want to know how a newly released model handles hard math before you cite it. You open MathArena, select ArXivLean June 2026, read the table, then click the Mistral Prover (Leanstral 1.5 + K3) cell to see the raw proof output behind its 38/46 verified solutions (82.61%).

Outcome: A specific, auditable number and the underlying model text to quote or reproduce.

Evaluation engineer

You are sanity-checking a benchmark claim from a vendor blog. You filter MathArena to the matching dataset month, compare the entry against Claude-Opus-5 (max) and Kimi K3, and open raw outputs to check whether scores came from real reasoning.

Outcome: A quick independent cross-check that either corroborates or contradicts the vendor's number.

Competition-math enthusiast

You follow the AIME and IMO crowd and want to see which model tops the charts this week. You browse the deprecated Final-Answer and Proof-Based views for context, then track the current dataset board to see which model added most recently.

Outcome: A clear, current read on bragging rights without wading through marketing pages.

Use Cases

Models Under the Hood

Mistral Prover (Leanstral 1.5 + K3)Muse Spark 1.3Muse Spark 1.1Claude-Opus-5 (max)Kimi K3

as of 2026-09-08

Limitations

  • MathArena is a free leaderboard focused exclusively on elite mathematical competitions, maintained by SRI Lab at ETH Zurich and INSAIT.
  • It covers dataset tracks (ArXivLean, BrokenArXiv, ArXivMath) plus deprecated views (Visual Math, Final-Answer Comps, Proof-Based Comps, Project Euler).
  • There is no mention of an API for automated benchmarking, so evaluation is manual and browser-based.
  • It is a free academic resource that may experience server slowdowns, and it offers no coding, chat, or multimodal scoring.

as of 2026-09-15

Verification history

We have re-verified Matharena 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. re-checked, vendor evidence unchanged
  5. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 7 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Matharena tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free

$0/mo

Ideal for

Researchers, evaluators, and math enthusiasts who want a free, citable read on LLM math performance.

What this tier adds

Starting tier and only tier — full leaderboard access, all dataset views, and clickable raw model outputs at $0.

Where the pricing makes sense

The company stage and team size where Matharena's pricing actually pencils out — and where peers do it cheaper.

MathArena is completely free — $0/mo for full leaderboard access, with no paid tiers, seat limits, or usage caps. That makes it cheaper than commercial eval platforms like Artificial Analysis or LMSYS-adjacent paid dashboards. The trade-off is scope: you get free math-competition rankings only, and no automation or API, so the real cost is your manual reading time.

Setup time & first value

How long it actually takes to get something useful out of Matharena — broken out by persona, not the marketing-page minute.

Zero setup: MathArena is a public website, so researchers get value in under a minute by picking a dataset month and reading the table. Evaluation engineers auditing outputs add a couple of minutes per cell to review raw model text. No install, no account, no API key.

Switching to or from Matharena

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • From manual paper benchmarking: Point your collaborators at the MathArena dataset views (ArXivLean, BrokenArXiv, ArXivMath) instead of re-running evals yourself for a first-pass number.
Migrating out
  • To an automated eval harness: MathArena has no API, so if you need CI-integrated benchmarking you will need to build your own benchmark runner outside it.

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Matharena”, and we withheld 6: 6 could not be judged, because “Matharena” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Matharena.

Official links

Tools that pair well with Matharena

Common stack mates teams adopt alongside Matharena, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Matharena

View all
Opencompass

Opencompass

Open-source LLM & VLM evaluation platform for standardized benchmarking

FreeTry
Goodfire

Goodfire

Silico: mechanistic interpretability platform to understand, debug, and design AI models

FreemiumTry
Fiddler AI

Fiddler AI

Fiddler AI is an enterprise AI control plane for agent observability, guardrails, and governance across the agentic lifecycle.

FreemiumTry

Frequently Asked Questions

Used Matharena? Help shape our editorial sentiment research.