LLM Stats

LLM Stats

Independent AI leaderboard scoring 400+ models from every major lab on one composite number that blends benchmark results with live API speed and pricing.

62/100MonitorFree planFreemium

For model selection, LLM Stats does the boring work most comparison sites skip: it keeps pricing and throughput in the same view as benchmark scores, and re-checks the pricing data every 30 minutes. The cost-per-turn table is the part I'd actually bring into a budget conversation — it shows Claude Fable 5.1 at $0.496 per turn and +2751% versus the median model, which a headline $/M token rate would hide. Pair it with Artificial Analysis if you want a second composite opinion, or with your own eval harness before committing spend.

Verified 5d ago · liveness 62/100 · cite: rightaichoice.com/tools/llm-stats

Best for
  • Developers shortlisting a model before running their own evals
  • Teams comparing cost per turn against benchmark score before committing spend
  • Researchers tracking benchmark movement across a 30d or 90d window
  • Buyers who need an open-weight alternative ranked against frontier models
Not ideal for
  • Teams that need output quality tested on their own proprietary data
  • Security, compliance or data-residency due diligence on a model vendor
  • Anyone wanting a ready-made AI assistant or chatbot product
Visit Website

IntermediateNo install and no account needed to read the leaderboard — you're looking at ranked models within a minute of landing, after a one-time human-verification click. The pricing hub and playground are equally direct. Programmatic use through the Data API or MCP server takes longer, since that's a client-side integration into your own tooling rather than a setup step on their side.WebAPI availableVerified 5d ago
Pricing
Free plan
FreemiumFree tier3 hidden costs
Learning curve
Intermediate
No install and no account needed to read the leaderboard — you're looking at ranked models within a minute of landing, after a one-time human-verification click. The pricing hub and playground are equally direct. Programmatic use through the Data API or MCP server takes longer, since that's a client-side integration into your own tooling rather than a setup step on their side.
Runs on
Web
API available · 1 integrations
Who it's for
Developer picking a default model for a coding agentProcurement lead modeling monthly inference spendResearch engineer tracking open-weight progress
Live sentiment
Is LLM Stats actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip LLM Stats if you need a model tested on your own proprietary data or a vendor's security and data-residency posture — this is a public-benchmark catalog, not an evaluation of how a model performs on your workload.

The 30-second take
Biggest gripe

Cost-per-turn figures are observations from the LLM Stats proxy, so a model that looks cheap on a $/M token basis can still land well above the median once verbosity is included — Claude Fable 5.1 shows +43% tokens and

Price reality

LLM Stats publishes a single Free tier at $0/mo covering the full 400-model leaderboard, category views, the Open LLM Leaderboard, comparison, pricing data, playground and newsletter. That makes it cheaper than paid model-evaluation gateways and free alongside analyst-style comparison services, provided you only need published benchmarks plus observed speed and price — not testing on your own data, where you'd need your own harness or an eval platform.

In short

LLM Stats — Independent AI leaderboard scoring 400+ models from every major lab on one composite number that blends benchmark results with live API speed and pricing. Best for Developers shortlisting a model before running their own evals, Teams comparing cost per turn against benchmark score before committing spend, Researchers tracking benchmark movement across a 30d or 90d window. Free to use.

What people actually say about LLM Stats — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

69 mentions across 4 sources (Hacker News, YouTube, Product Hunt, Lemmy) · researched Jul 3, 2026.

40% positive60% critical

Average across the 4 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Aggregates 300+ models with one composite score for quick comparison.
  • +Side-by-side cost-per-token next to benchmark scores saves time.
  • +Playground lets you test models live before committing to an API.
  • +Filters by use case (coding, writing, math) directly address buyer needs.
  • +Open LLM leaderboard helps compare open-weight alternatives fairly.
Recurring frustrations
  • −Update frequency is unclear, worrying users about stale data.
  • −No integrations with tools like Raycast, limiting workflow use.
  • −Third-party model providers may introduce latency or pricing gaps.
  • −Support responsiveness is unknown due to limited community feedback.
  • −Data sources behind benchmarks aren't fully transparent.
Patterns worth knowing
Great for quick model comparison, but data freshness is a concern
Seen on Product Hunt, Hacker News
Free tier is generous and useful for exploratory research
Seen on Product Hunt
Playground and chat features are valued for hands-on testing
Seen on Product Hunt
Learning curve
beginnerProductive in ~5 minutes
Hidden costs people mention
  • • Playground credits may run out on free tier requiring upgrade
  • • API usage beyond fair use may incur additional fees

Viability Score

62/100
Monitor

How well maintained and how widely used is LLM Stats? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
40
What the vendor publishes
0

Last calculated: October 2026

How we score →

Key Features

  • Composite LLM Stats Score blending benchmarks, speed and price
  • Leaderboard ranking 400 models across all major labs
  • Category leaderboards for Coding, Writing, Math, Research and Reasoning
  • Dedicated Long Context and Tool Calling rankings
  • Open LLM Leaderboard filtered to publicly released weights
  • Side-by-side model comparison on price, context, speed and benchmarks
  • Observed cost-per-turn comparison across 96 models vs the median model
  • Per-token input, cached input and output pricing across 161 priced models
  • Pricing data refreshed every 30 minutes with billing-sample cross-checks
  • Output speed tracking in tokens per second (up to 660 tok/s listed)
  • Context window comparison up to 2.0M tokens
  • New models feed covering releases from the last 15 days
  • Community arenas for Chat, Coding, Image and Video evaluation
  • Playground for interacting with hundreds of models
  • Benchmark library covering GPQA, MMLU, SWE-Bench, AIME and LiveCodeBench

About LLM Stats

FreemiumIntermediateAPI availableWeb

LLM Stats is an independent AI leaderboard and model comparison hub. Its homepage table currently ranks 400 models — including Claude Opus 5.5, GPT-6 Astra, Claude Sonnet 5.5, Kimi K3, Gemini 3.1 Pro and DeepSeek-V4.1-Flash — with one composite LLM Stats Score that blends published benchmark results (GPQA Diamond, SWE-Bench Verified, MMLU, AIME, coding-arena play) with live API metrics like tokens-per-second output speed and per-token pricing. That lets a frontier proprietary model and a cheap open-weight challenger sit on the same axis, which is the thing most vendor benchmark tables won't do. It's built for developers picking a model to ship on, researchers tracking benchmark movement over a 30d or 90d window, and procurement teams weighing cost per usable token. The full leaderboard filters by time window, price, parameters, license and model type, and there are dedicated views for Coding, Writing, Math, Research, Long Context, Tool Calling, Reasoning, Image Gen, Video Creation, TTS, STT and Embeddings. A separate Open LLM Leaderboard isolates open-weight checkpoints, and the compare view puts two models side by side on pricing, context, throughput and benchmark scores. The pricing hub goes past a rate card: it reports observed cost per turn across 96 observed models against the median model, which is the number that actually shows up on your bill. Freshness is the pitch. The leaderboard refreshes as new benchmark results land, the pricing page re-checks data every 30 minutes, and a New Models feed surfaces releases from the past 15 days. A playground lets you poke at hundreds of models directly, weekly newsletter digests filter the noise, and a Data API plus MCP server expose the catalog programmatically. Access requires a quick human-verification step on first visit.

Behind the Verdict

The core value here is the axis problem. Benchmarks tell you a model is smart; they don't tell you it costs +2751% more per turn than the median model because it's four times as verbose. LLM Stats puts both on one screen. The pricing page lists 377 models with 161 priced and shows observed cost per turn for 96 of them, including how each model's token count compares to the median — that second number is what turns a rate card into a real forecast. Strengths worth calling out. The category leaderboards go deeper than one general ranking: Coding, Writing, Math, Research, Long Context, Tool Calling, Reasoning, Image Gen, Video Creation, TTS, STT and Embeddings are all separate views, so a TTS buyer isn't sifting through reasoning scores. The Open LLM Leaderboard isolates models with publicly released weights — Kimi K3 currently leads that group on the homepage summary at 93.5% GPQA and it sits 10th overall at $3.39/M tokens, which is the kind of side-by-side a self-hosting team needs. The compare view handles pricing, context length, throughput and benchmarks in one place, and the leaderboard filters by 30d/90d window so you can spot a model that's improving versus one coasting on an old score. Weaknesses are mostly about what a composite score can't tell you. The methodology page describes scores as uncertainty-aware aggregates, not truth — the homepage FAQ itself says "best" depends on what you're optimizing for and names a different winner per axis (GPT-6 Astra on GPQA reasoning, Claude Opus 5 on coding-arena play, Gemini 3.1 Pro as cheapest in the top 10 at $2.00/M tok). None of that predicts how a model handles your proprietary data, your prompt style, or your latency budget in a specific region. Speed is measured across supported provider APIs over a recent window, so it's an average, not your p99. And there's a human-verification gate on first visit, which is friction if you just want to glance at a number. Where it fits: as a shortlisting and cost-modeling layer before you run your own evals, and as a weekly tracking habit if you care about release cadence across OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, Alibaba and Zhipu. Where it doesn't: it is a catalog, not a provider. There's no fine-tuning, hosting or inference here, and nothing resembling a ready-made assistant. If you need security, compliance or data-residency answers about a vendor, this is not the source.

Researching LLM Stats? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas LLM Stats actually fits — and what changes day-one when you adopt it.

Developer picking a default model for a coding agent

Open the Coding leaderboard, filter to the 90d window, then use compare to put Claude Opus 5 against Kimi K3 on benchmark score, context window, output speed and per-token price.

Outcome: A two-model shortlist with a defensible cost estimate per turn, ready to validate against your own task suite.

Procurement lead modeling monthly inference spend

Open the pricing hub, sort by observed cost per turn, and pull the models already in your stack against the median-model baseline.

Outcome: A spend forecast that accounts for verbosity, not just the $/M token headline — the +2751% vs median figure on Claude Fable 5.1 is the kind of line item a rate card hides.

Research engineer tracking open-weight progress

Use the Open LLM Leaderboard to check where Kimi K3, GLM-5.3 and DeepSeek-V4.1-Flash sit against frontier proprietary models, then set up the New Models feed for the last 15 days.

Outcome: A weekly read on whether self-hosting is competitive yet for a given task, without manually reconciling three labs' blog posts.

Use Cases

  • Compare Claude Opus 5.5, GPT-6 Astra and Gemini 3.1 Pro on reasoning, coding and cost before you pick a default model.
  • Model monthly spend using observed cost-per-turn rather than a headline $/M token rate.
  • Filter the leaderboard to a 30-day window to see which models gained or slipped recently.
  • Find an open-weight alternative — Kimi K3, GLM-5.3, DeepSeek-V4.1-Flash — ranked against frontier proprietary models.
  • Check output throughput before committing to a model for a streaming chat UI or an agentic loop.
  • Verify the largest usable context window when you need long-document or long-trace workloads.
  • Test the same prompt across multiple models in the playground before writing your own harness.
  • Pull rankings programmatically through the Data API or MCP server for an internal cost dashboard.

Models Under the Hood

Claude Opus 5.5GPT-6 AstraClaude Opus 5Claude Fable 5.1GPT-5.6 SolClaude Mythos PreviewClaude Fable 5Muse Spark 1.3Kimi K3GLM-5.3Atria Dawn PreviewQwen3.8 Max

as of 2026-09-25

Limitations

  • The rankings rely on public benchmarks and live API metrics, so they may not predict real-world performance on specialized or proprietary tasks.
  • The composite score is an aggregate — the homepage FAQ itself points out the leader model differs per axis, so a single number is a starting filter, not a verdict.
  • Output speed is summarized over a recent observation window across supported provider APIs, which means it's an average rather than peak or per-region latency.
  • Context windows are advertised figures; the site notes providers vary in how well they use the upper end, and directs you to per-model effective-context notes.
  • First visit requires a human-verification step, which adds friction to a quick lookup.

as of 2026-10-03

Verification history

We have re-verified LLM Stats 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-checked, vendor evidence unchanged
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 8 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published LLM Stats tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free

$0/mo

Ideal for

Developers, researchers and procurement teams who need the full 400-model leaderboard, pricing data and playground without a subscription.

What this tier adds

Starting tier — $0/mo and everything is included: all leaderboards, category views, comparison, cost-per-turn data, playground, newsletter, Data API and MCP server. Access requires a one-time human-verification step.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Cost-per-turn figures are observations from the LLM Stats proxy, so a model that looks cheap on a $/M token basis can still land well above the median once verbosity is included — Claude Fable 5.1 shows +43% tokens and
  • Coverage is uneven: the pricing hub lists 377 models but only 161 are priced and only 96 have observed cost-per-turn data, so a model you're evaluating may have no cost number at all.
  • Free access runs behind a human-verification step on first visit, which costs time if you're pulling figures during a live vendor negotiation.

Where the pricing makes sense

The company stage and team size where LLM Stats's pricing actually pencils out — and where peers do it cheaper.

LLM Stats publishes a single Free tier at $0/mo covering the full 400-model leaderboard, category views, the Open LLM Leaderboard, comparison, pricing data, playground and newsletter. That makes it cheaper than paid model-evaluation gateways and free alongside analyst-style comparison services, provided you only need published benchmarks plus observed speed and price — not testing on your own data, where you'd need your own harness or an eval platform.

Setup time & first value

How long it actually takes to get something useful out of LLM Stats — broken out by persona, not the marketing-page minute.

No install and no account needed to read the leaderboard — you're looking at ranked models within a minute of landing, after a one-time human-verification click. The pricing hub and playground are equally direct. Programmatic use through the Data API or MCP server takes longer, since that's a client-side integration into your own tooling rather than a setup step on their side.

Switching to or from LLM Stats

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a vendor benchmark page: look up the same model on the LLM Stats leaderboard to see it scored against 399 competitors rather than only its own family.
  • →From a bookmarked static leaderboard: switch to the 30d/90d window filters so rankings reflect recent benchmark results rather than a page last edited months ago.
  • →From spreadsheet cost modeling: replace your $/M token column with the observed cost-per-turn table, which accounts for how verbose each model actually is.
Migrating out
  • ↗To your own eval harness: export the shortlist and rerun the same prompts on your proprietary data, since public benchmarks can't answer task-specific quality.
  • ↗To a vendor's published rate card: go direct once you've picked a model and need contractual and compliance detail LLM Stats doesn't cover.
  • ↗To a second aggregator: cross-check the composite score against another independent ranking before treating any single number as final.

Integrations

MCP

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “LLM Stats”, and we withheld 6: 6 did not mention LLM Stats. We are showing none, because we could not prove any of them are about LLM Stats.

Tools that pair well with LLM Stats

Common stack mates teams adopt alongside LLM Stats, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to LLM Stats

View all
Arena AI

Arena AI

Arena AI is a free, community-voted LLM leaderboard ranking chat models, agents, and fullstack code on live head-to-head battles.

FreemiumTry
Opencompass

Opencompass

OpenCompass (司南) benchmarks LLMs, VLMs, and AI4S models against 100+ open evaluation datasets with published, dated leaderboards

FreeTry
ClawBench

ClawBench

ClawBench benchmarks AI browser agents on live websites with HTTP-interception scoring and LLM-judge grading.

FreeTry

Frequently Asked Questions

Used LLM Stats? Help shape our editorial sentiment research.