ClawBench

ClawBench

Open-source benchmark that runs AI browser agents on live websites and scores the HTTP request they actually fire.

70/100Safe BetFreeFree

If you build or buy browser agents, ClawBench gives you the number most vendors avoid: what happens when the agent touches a real checkout page. The two-stage split — deterministic HTTP interception, then a published judge rubric — means you can re-run any cell with your own judge and see if the score holds. The gap between claude-opus-4-7's 44.6% Reward and its 54.6% Intercepted is the honest headline: agents fire the right request far more often than they satisfy the instruction. Compare that with claude-opus-4-6 on the V1 board at 61.4%, which is a different corpus draw rather than a model regression. Pair it with WebArena or Mind2Web when you want controlled sandboxes, and treat

Verified 13h ago · liveness 70/100 · cite: rightaichoice.com/tools/clawbench

Best for
  • AI researchers benchmarking browser agents on live sites instead of frozen snapshots
  • Agent developers who need failure-mode diagnosis from synchronized video, HAR and LLM turns
  • Model evaluation teams comparing cost-per-task alongside pass rates
  • Training teams mining JSONL trajectories for SFT, DPO or PRM datasets
Not ideal for
  • Casual users wanting a ready-to-use browser agent rather than an evaluation harness
  • Teams without CLI or Python skills who need a click-to-run evaluation UI
  • Air-gapped environments, since scoring depends on reaching live third-party websites
Visit Website

AdvancedResearchers with Python experience: roughly 15-30 minutes to pip install clawbench-eval, pull a corpus and complete a first run against a model. Agent developers wiring in their own harness: a few hours to a day, since you supply the browser, driver and agent loop around the evaluator. Teams expecting a hosted UI: no path — it's a CLI plus a trace browser.Web · CLINo public APIVerified 13h ago
Pricing
Free
FreeFree tier4 hidden costs
Learning curve
Advanced
Researchers with Python experience: roughly 15-30 minutes to pip install clawbench-eval, pull a corpus and complete a first run against a model. Agent developers wiring in their own harness: a few hours to a day, since you supply the browser, driver and agent loop around the evaluator. Teams expecting a hosted UI: no path — it's a CLI plus a trace browser.
Runs on
WebCLI
No public API · 5 integrations
Who it's for
Agent researcherFailure-mode debuggerTraining-data engineer
Live sentiment
Is ClawBench actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip ClawBench if you want a managed, click-to-run evaluation service or you're ranking models on two-point gaps — 283 tasks is a thin sample at that resolution.

The 30-second take
Biggest gripe

Running the corpus against frontier models is not free at the model end: on the V2 board, claude-opus-4-7 costs $4.4425 per task against $0.0721 for deepseek-v4-pro, so a full 130-task sweep differs by roughly $568

Price reality

ClawBench itself is an Apache-2.0 open-source benchmark with the full corpus, leaderboard and trace bundles published; your real spend is model inference during evaluation runs. That makes it cheaper than commercial eval vendors who charge per seat or per run, but more expensive than a pure sandbox benchmark like WebArena or Mind2Web, where tasks are frozen and no live-site inference or navigation cost is incurred.

In short

ClawBench — Open-source benchmark that runs AI browser agents on live websites and scores the HTTP request they actually fire. Best for AI researchers benchmarking browser agents on live sites instead of frozen snapshots, Agent developers who need failure-mode diagnosis from synchronized video, HAR and LLM turns, Model evaluation teams comparing cost-per-task alongside pass rates. Free to use.

What's new in ClawBench

Checked today

Across the latest 1 update: 1 changelog entry.

What people actually say about ClawBench — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

29 mentions across 4 sources (YouTube, Bluesky, GitHub, Lemmy) · researched Jul 6, 2026.

61% positive39% critical

Average across the 4 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Two-stage scoring (HTTP interception + LLM judge) adds honesty.
  • +130+ real live tasks across diverse platforms.
  • +Public leaderboard with cost/task metrics for model comparison.
  • +Open-source dataset and traces on Hugging Face.
  • +CLI tool (pip install clawbench-eval) for easy local runs.
Recurring frustrations
  • −Very little community discussion to validate ease of use.
  • −Some test cases have instruction conflicts and placeholder bugs.
  • −No pre-built Docker containers; must build locally.
  • −Running local models requires expensive hardware.
  • −Documentation on recommended model/harness pairings is unclear.
Patterns worth knowing
ClawBench is a valuable but niche benchmark for AI agent research, praised for its realistic live tasks and honest scoring.
Seen on Bluesky, YouTube, Lemmy
Users want better onboarding: pre-built containers, clearer model-harness guidance, and resolved test case bugs.
Seen on GitHub
Hardware cost and speed for local model evaluation is a concern, limiting accessibility for hobbyists.
Seen on YouTube, Lemmy
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • Requires cloud API keys or local GPU hardware for model inference
  • • Computational cost of running LLM judge and tasks adds up

Viability Score

70/100
Safe Bet

How well maintained and how widely used is ClawBench? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
61
What the vendor publishes
20

Last calculated: October 2026

How we score →

Key Features

  • Live-website agent testing on real platforms instead of static snapshots or sandboxes
  • Stage 1 deterministic HTTP interception: final request URL and method checked against the task schema
  • Stage 2 LLM judge (deepseek/deepseek-v4-pro) reads the intercepted payload against the instruction
  • Dual rubrics: lenient Reward (no contradiction → match) and Reward (strict) (ambiguous → mismatch)
  • 283 distinct everyday tasks across 163 live platforms as of the 2026-05-20 snapshot
  • 1,724 judge-verified runs spanning 13 frontier models
  • Cost-per-task and Pass/Total columns on every leaderboard row
  • 5-layer trace bundle: video, actions, agent messages, HTTP requests and graded verdict
  • Six time-synchronized signals per run on one clock (~80 events, ~150 LLM turns, ~500 HTTP calls)
  • Replay and audit: step through video, HAR and agent reasoning side-by-side
  • CLI evaluation: pip install clawbench-eval && clawbench run --corpus v2 --model your-model
  • JSONL-native trajectories ready for SFT, DPO and PRM fine-tuning
  • 918 V1 + 806 V2 frontier-model trajectories for training and failure-pair mining
  • Searchable, filterable task gallery with no download required
  • Adapters for the claw-eval, WildClawBench and ClawMark corpora on the same evaluator

About ClawBench

FreeAdvancedNo APIWeb · CLI

ClawBench is an open-source benchmark that gives an AI agent a plain everyday instruction — order one Pad Thai on Uber Eats with a 'no peanuts' note, book a flight, apply for a job — and lets it drive a real browser on the real site. Nothing is a frozen snapshot: the harness waits for the agent's final submit request and intercepts it before it fires. Scoring runs in two stages. Stage 1 (Intercepted) is deterministic: did the final request's URL and method match the task schema? Stage 2 (Reward) passes the intercepted payload to deepseek/deepseek-v4-pro, which reads it against the instruction under a lenient 'no contradiction → match' rubric; Reward (strict) applies 'ambiguous → mismatch' instead. The leaderboard ranks by lenient Reward. The 2026-05-20 snapshot reports 1,724 judge-verified runs, 13 frontier models, 283 distinct tasks and 163 live platforms, refreshed weekly. On the V2 (Hermes) board claude-opus-4-7 leads at 44.6% Reward and 54.6% Intercepted (58/130), then gpt-5.5 at 35.4%, glm-5.1 at 34.6% and deepseek-v4-pro at 33.9% — a spread wide enough to separate models rather than tie them. Every run ships as a five-layer trace bundle — video, actions, agent messages, HTTP requests and the graded verdict — plus run metadata, all on one clock, with roughly 80 events, ~150 LLM turns and ~500 HTTP calls per task. Click frame 1872 of recording.mp4 and you can find the exact event and LLM turn that fired the request. The dataset is Apache-2.0 on GitHub and Hugging Face with cross-org mirrors (NAIL-Group, TIGER-Lab), installs via pip (clawbench-eval), and exports JSONL trajectories: 918 V1 plus 806 V2 frontier-model runs, SFT/DPO/PRM-ready. It is built for AI researchers, agent developers and model-eval teams who want replayable numbers instead of a hosted score. Sandbox benchmarks like WebArena and Mind2Web trade realism for control; ClawBench goes the other way, accepting live-site volatility in exchange for tasks that can't be memorized.

Behind the Verdict

ClawBench's core bet is that agent evaluation should happen where the money moves: on live third-party sites, against real checkout and application flows. That bet buys you the two-stage design. Stage 1 is deterministic — the harness intercepts the agent's final HTTP request and checks URL and method against the task schema, so a run only counts if something real was about to fire. Stage 2 hands the intercepted payload to deepseek/deepseek-v4-pro, which compares it against the instruction with a lenient rubric ('no contradiction → match') for the headline Reward and a strict rubric ('ambiguous → mismatch') for Reward (strict). Because both rubrics and the judge are published, you can re-run any cell — your judge on their data, or theirs on your judge — which is a level of auditability most leaderboards don't offer. The second strength is the trace bundle. Each run ships six time-synchronized signals on one clock with roughly 80 events, ~150 LLM turns and ~500 HTTP calls per task. Video, actions.jsonl, agent-messages.jsonl, requests.jsonl, interception.json and run-meta.json sit side by side, so clicking frame 1872 of recording.mp4 lets you find the exact action event, the LLM turn that triggered it, and the HTTP request it fired. For diagnosing why an agent failed — did it pick the wrong item, miss the note, or never submit? — that is more useful than a pass rate. The training angle is real too: 918 V1 plus 806 V2 frontier-model trajectories, JSONL-native and SFT/DPO/PRM-ready, across 13 models on identical tasks, so success-versus-failure pairs are aligned by task rather than guessed at. The honest downsides. Reward carries judge variability that Intercepted does not — the headline number is a model's opinion, and the strict column shows how much that opinion moves (claude-opus-4-7 falls from 44.6% to 24.6%). Denominators shift between boards: V2 (Hermes) tracks 130 tasks across 63 live platforms, while the dataset-wide figures are 283 tasks and 163 platforms, so V1 and V2 tables aren't eyeball-comparable. The V2 (OpenClaw) table lists a single model at 0.0%, and V2 (Codex) and V2 (Claude Code) are empty at this snapshot. And 283 tasks is a thin sample if you're trying to separate models on a two-point gap. Where it fits: agent research teams, eval groups who need cost-per-task next to pass rate, training teams mining trajectories. Where it doesn't: teams that want a click-to-run hosted eval UI, non-technical buyers, and air-gapped environments, since live-site scoring needs outbound access to third-party sites.

Researching ClawBench? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas ClawBench actually fits — and what changes day-one when you adopt it.

Agent researcher

You install clawbench-eval via pip and run clawbench run --corpus v2 --model your-model to place your agent on the V2 (Hermes) board alongside claude-opus-4-7 and gpt-5.5.

Outcome: You get Intercepted, Reward, Reward (strict), cost per task and Pass/Total (e.g. 58/130) on the same row, plus a downloadable 5-layer trace bundle per run.

Failure-mode debugger

An agent passes Intercepted at 54.6% but only reaches 44.6% Reward. You open the trace browser and step through recording.mp4, actions.jsonl, agent-messages.jsonl and requests.jsonl on one clock.

Outcome: You can click a video frame, find the exact action event and LLM turn behind it, and pinpoint whether the agent picked the wrong item, dropped the 'no peanuts' note, or submitted the wrong address.

Training-data engineer

You download the 918 V1 and 806 V2 frontier-model trajectories and mine success-versus-failure pairs across 13 models on identical tasks.

Outcome: You have JSONL-native data ready for SFT, DPO or PRM without building your own task corpus or alignment pass.

Use Cases

  • Evaluate a browser agent on live tasks like booking flights, ordering food, or applying for jobs.
  • Compare agent performance with deterministic HTTP interception plus LLM-based reward scoring.
  • Submit your model to the public leaderboard and track cost per task against 13 frontier models.
  • Download 5-layer trace bundles to diagnose failure modes and audit judge calls.
  • Run ClawBench or scope peers with a single CLI command.
  • Fine-tune your own agent on 918 V1 + 806 V2 JSONL trajectories with SFT/DPO/PRM.
  • Reproduce a leaderboard cell by re-running our judge on your data or your judge on ours.
  • Browse all 283 task definitions in a searchable, filterable gallery without downloading.

Models Under the Hood

claude-opus-4-7claude-opus-4-6Claude Sonnet 4.6claude-haiku-4-5-20251001gpt-5.5gpt-5.4-2026-03-05gpt-5.4-mini-2026-03-17glm-5.1deepseek-v4-prodeepseek-v4-flash:free

as of 2026-10-10

Limitations

  • ClawBench is a CLI-first benchmark: setup means pip install clawbench-eval and commands like clawbench run --corpus v2 --model your-model.
  • The V2 (Hermes) board tracks 130 tasks across 63 live platforms, while the dataset-wide figures are 283 tasks and 163 platforms across all corpora — the differing denominators make eyeball comparisons between V1 and V2 tables easy to misread.
  • Stage 2 scoring relies on an LLM judge (deepseek/deepseek-v4-pro) with lenient and strict rubrics, so the headline Reward number carries judgment variability that deterministic interception alone doesn't; claude-opus-4-7 drops from 44.6% Reward to 24.6% under the strict rubric.
  • Leaderboard coverage thins out in places: the V2 (OpenClaw) table has a single model at 0.0%, and V2 (Codex) and V2 (Claude Code) are empty at this snapshot.
  • Weekly refresh means a pinned number is only current as of its snapshot date (last 2026-05-20, when the corpus stood at 283 tasks and 163 live platforms).
  • Because tasks run against live third-party sites, results are volatile by design and old traces won't reproduce byte-for-byte.

as of 2026-10-10

Verification history

We have re-verified ClawBench 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-checked, vendor evidence unchanged
  3. — re-checked, vendor evidence unchanged
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-checked, vendor evidence unchanged
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 8 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published ClawBench tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free

$0

Ideal for

AI researchers, agent developers and eval teams who need live-site browser-agent numbers they can re-run and audit, not a hosted score.

What this tier adds

Starting tier — the whole benchmark: Apache-2.0 corpus, clawbench-eval CLI, 1,724 runs, 5-layer trace bundles and 918 V1 + 806 V2 JSONL trajectories.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Running the corpus against frontier models is not free at the model end: on the V2 board, claude-opus-4-7 costs $4.4425 per task against $0.0721 for deepseek-v4-pro, so a full 130-task sweep differs by roughly $568
  • Scoring assumes outbound access to live third-party sites, so an air-gapped research cluster needs a network path before any run works.
  • Self-hosting means maintaining your own agent harness, browser and driver stack — the benchmark is Apache-2.0, but the environment around it is yours to keep running.
  • Weekly snapshots replace earlier numbers, so a leaderboard cell you cite in a paper can be stale within seven days unless you pin the snapshot date.

Where the pricing makes sense

The company stage and team size where ClawBench's pricing actually pencils out — and where peers do it cheaper.

ClawBench itself is an Apache-2.0 open-source benchmark with the full corpus, leaderboard and trace bundles published; your real spend is model inference during evaluation runs. That makes it cheaper than commercial eval vendors who charge per seat or per run, but more expensive than a pure sandbox benchmark like WebArena or Mind2Web, where tasks are frozen and no live-site inference or navigation cost is incurred.

Setup time & first value

How long it actually takes to get something useful out of ClawBench — broken out by persona, not the marketing-page minute.

Researchers with Python experience: roughly 15-30 minutes to pip install clawbench-eval, pull a corpus and complete a first run against a model. Agent developers wiring in their own harness: a few hours to a day, since you supply the browser, driver and agent loop around the evaluator. Teams expecting a hosted UI: no path — it's a CLI plus a trace browser.

Switching to or from ClawBench

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From WebArena: port your WebArena-style task definitions onto the ClawBench evaluator, then re-point the same agent loop at live sites and compare Intercepted and Reward against the published V2 (Hermes) board.
  • →From Mind2Web: keep your frozen-snapshot pass rates as a baseline and add ClawBench runs on the 130 V2 tasks to see which of your gains survive live-site volatility.
  • →From claw-eval / WildClawBench / ClawMark: reuse the existing adapters to score your runs on the ClawBench evaluator without rewriting the harness.
Migrating out
  • ↗To WebArena: freeze the ClawBench task list into a sandbox if you need reproducible, byte-identical runs and can accept memorizable tasks.
  • ↗To Mind2Web: move to static snapshot evaluation when you need a stable denominator and don't care about live checkout or application flows.
  • ↗To WebVoyager-style suites: pick a smaller public suite when 130 tasks and a two-rubric LLM judge is more machinery than your question needs.

Integrations

GitHubHugging FacePyPIGradioOpenRouter

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “ClawBench”, and we withheld 6: 6 could not be judged, because “ClawBench” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about ClawBench.

Official links

Tools that pair well with ClawBench

Common stack mates teams adopt alongside ClawBench, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to ClawBench

View all
ClawX

ClawX

Free, open-source desktop AI assistant that runs 24/7 on your own machine, scraping and analyzing web sources autonomously.

FreemiumTry
Steel Browser

Steel Browser

Open-source cloud browser API for running fleets of AI agent browser sessions at scale

FreemiumTry
Arena AI

Arena AI

A free public AI leaderboard where live head-to-head votes and guided tasks rank chat models, coding agents, and fullstack builds.

FreemiumTry

Frequently Asked Questions

Used ClawBench? Help shape our editorial sentiment research.