ClawBench
Open-source benchmark that runs AI browser agents on live websites and scores the HTTP request they actually fire.
If you build or buy browser agents, ClawBench gives you the number most vendors avoid: what happens when the agent touches a real checkout page. The two-stage split — deterministic HTTP interception, then a published judge rubric — means you can re-run any cell with your own judge and see if the score holds. The gap between claude-opus-4-7's 44.6% Reward and its 54.6% Intercepted is the honest headline: agents fire the right request far more often than they satisfy the instruction. Compare that with claude-opus-4-6 on the V1 board at 61.4%, which is a different corpus draw rather than a model regression. Pair it with WebArena or Mind2Web when you want controlled sandboxes, and treat
Verified 13h ago · liveness 70/100 · cite: rightaichoice.com/tools/clawbench
- AI researchers benchmarking browser agents on live sites instead of frozen snapshots
- Agent developers who need failure-mode diagnosis from synchronized video, HAR and LLM turns
- Model evaluation teams comparing cost-per-task alongside pass rates
- Training teams mining JSONL trajectories for SFT, DPO or PRM datasets
- Casual users wanting a ready-to-use browser agent rather than an evaluation harness
- Teams without CLI or Python skills who need a click-to-run evaluation UI
- Air-gapped environments, since scoring depends on reaching live third-party websites
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip ClawBench if you want a managed, click-to-run evaluation service or you're ranking models on two-point gaps — 283 tasks is a thin sample at that resolution.
Running the corpus against frontier models is not free at the model end: on the V2 board, claude-opus-4-7 costs $4.4425 per task against $0.0721 for deepseek-v4-pro, so a full 130-task sweep differs by roughly $568
ClawBench itself is an Apache-2.0 open-source benchmark with the full corpus, leaderboard and trace bundles published; your real spend is model inference during evaluation runs. That makes it cheaper than commercial eval vendors who charge per seat or per run, but more expensive than a pure sandbox benchmark like WebArena or Mind2Web, where tasks are frozen and no live-site inference or navigation cost is incurred.
In short
ClawBench — Open-source benchmark that runs AI browser agents on live websites and scores the HTTP request they actually fire. Best for AI researchers benchmarking browser agents on live sites instead of frozen snapshots, Agent developers who need failure-mode diagnosis from synchronized video, HAR and LLM turns, Model evaluation teams comparing cost-per-task alongside pass rates. Free to use.
What's new in ClawBench
Checked todayAcross the latest 1 update: 1 changelog entry.
What people actually say about ClawBench — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
29 mentions across 4 sources (YouTube, Bluesky, GitHub, Lemmy) · researched Jul 6, 2026.
Average across the 4 sources that answered — each source counts once, not each post.
- +Two-stage scoring (HTTP interception + LLM judge) adds honesty.
- +130+ real live tasks across diverse platforms.
- +Public leaderboard with cost/task metrics for model comparison.
- +Open-source dataset and traces on Hugging Face.
- +CLI tool (pip install clawbench-eval) for easy local runs.
- −Very little community discussion to validate ease of use.
- −Some test cases have instruction conflicts and placeholder bugs.
- −No pre-built Docker containers; must build locally.
- −Running local models requires expensive hardware.
- −Documentation on recommended model/harness pairings is unclear.
- • Requires cloud API keys or local GPU hardware for model inference
- • Computational cost of running LLM judge and tasks adds up
Viability Score
How well maintained and how widely used is ClawBench? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Live-website agent testing on real platforms instead of static snapshots or sandboxes
- Stage 1 deterministic HTTP interception: final request URL and method checked against the task schema
- Stage 2 LLM judge (deepseek/deepseek-v4-pro) reads the intercepted payload against the instruction
- Dual rubrics: lenient Reward (no contradiction → match) and Reward (strict) (ambiguous → mismatch)
- 283 distinct everyday tasks across 163 live platforms as of the 2026-05-20 snapshot
- 1,724 judge-verified runs spanning 13 frontier models
- Cost-per-task and Pass/Total columns on every leaderboard row
- 5-layer trace bundle: video, actions, agent messages, HTTP requests and graded verdict
- Six time-synchronized signals per run on one clock (~80 events, ~150 LLM turns, ~500 HTTP calls)
- Replay and audit: step through video, HAR and agent reasoning side-by-side
- CLI evaluation: pip install clawbench-eval && clawbench run --corpus v2 --model your-model
- JSONL-native trajectories ready for SFT, DPO and PRM fine-tuning
- 918 V1 + 806 V2 frontier-model trajectories for training and failure-pair mining
- Searchable, filterable task gallery with no download required
- Adapters for the claw-eval, WildClawBench and ClawMark corpora on the same evaluator
About ClawBench
ClawBench is an open-source benchmark that gives an AI agent a plain everyday instruction — order one Pad Thai on Uber Eats with a 'no peanuts' note, book a flight, apply for a job — and lets it drive a real browser on the real site. Nothing is a frozen snapshot: the harness waits for the agent's final submit request and intercepts it before it fires. Scoring runs in two stages. Stage 1 (Intercepted) is deterministic: did the final request's URL and method match the task schema? Stage 2 (Reward) passes the intercepted payload to deepseek/deepseek-v4-pro, which reads it against the instruction under a lenient 'no contradiction → match' rubric; Reward (strict) applies 'ambiguous → mismatch' instead. The leaderboard ranks by lenient Reward. The 2026-05-20 snapshot reports 1,724 judge-verified runs, 13 frontier models, 283 distinct tasks and 163 live platforms, refreshed weekly. On the V2 (Hermes) board claude-opus-4-7 leads at 44.6% Reward and 54.6% Intercepted (58/130), then gpt-5.5 at 35.4%, glm-5.1 at 34.6% and deepseek-v4-pro at 33.9% — a spread wide enough to separate models rather than tie them. Every run ships as a five-layer trace bundle — video, actions, agent messages, HTTP requests and the graded verdict — plus run metadata, all on one clock, with roughly 80 events, ~150 LLM turns and ~500 HTTP calls per task. Click frame 1872 of recording.mp4 and you can find the exact event and LLM turn that fired the request. The dataset is Apache-2.0 on GitHub and Hugging Face with cross-org mirrors (NAIL-Group, TIGER-Lab), installs via pip (clawbench-eval), and exports JSONL trajectories: 918 V1 plus 806 V2 frontier-model runs, SFT/DPO/PRM-ready. It is built for AI researchers, agent developers and model-eval teams who want replayable numbers instead of a hosted score. Sandbox benchmarks like WebArena and Mind2Web trade realism for control; ClawBench goes the other way, accepting live-site volatility in exchange for tasks that can't be memorized.
Behind the Verdict
ClawBench's core bet is that agent evaluation should happen where the money moves: on live third-party sites, against real checkout and application flows. That bet buys you the two-stage design. Stage 1 is deterministic — the harness intercepts the agent's final HTTP request and checks URL and method against the task schema, so a run only counts if something real was about to fire. Stage 2 hands the intercepted payload to deepseek/deepseek-v4-pro, which compares it against the instruction with a lenient rubric ('no contradiction → match') for the headline Reward and a strict rubric ('ambiguous → mismatch') for Reward (strict). Because both rubrics and the judge are published, you can re-run any cell — your judge on their data, or theirs on your judge — which is a level of auditability most leaderboards don't offer. The second strength is the trace bundle. Each run ships six time-synchronized signals on one clock with roughly 80 events, ~150 LLM turns and ~500 HTTP calls per task. Video, actions.jsonl, agent-messages.jsonl, requests.jsonl, interception.json and run-meta.json sit side by side, so clicking frame 1872 of recording.mp4 lets you find the exact action event, the LLM turn that triggered it, and the HTTP request it fired. For diagnosing why an agent failed — did it pick the wrong item, miss the note, or never submit? — that is more useful than a pass rate. The training angle is real too: 918 V1 plus 806 V2 frontier-model trajectories, JSONL-native and SFT/DPO/PRM-ready, across 13 models on identical tasks, so success-versus-failure pairs are aligned by task rather than guessed at. The honest downsides. Reward carries judge variability that Intercepted does not — the headline number is a model's opinion, and the strict column shows how much that opinion moves (claude-opus-4-7 falls from 44.6% to 24.6%). Denominators shift between boards: V2 (Hermes) tracks 130 tasks across 63 live platforms, while the dataset-wide figures are 283 tasks and 163 platforms, so V1 and V2 tables aren't eyeball-comparable. The V2 (OpenClaw) table lists a single model at 0.0%, and V2 (Codex) and V2 (Claude Code) are empty at this snapshot. And 283 tasks is a thin sample if you're trying to separate models on a two-point gap. Where it fits: agent research teams, eval groups who need cost-per-task next to pass rate, training teams mining trajectories. Where it doesn't: teams that want a click-to-run hosted eval UI, non-technical buyers, and air-gapped environments, since live-site scoring needs outbound access to third-party sites.
Researching ClawBench? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas ClawBench actually fits — and what changes day-one when you adopt it.
You install clawbench-eval via pip and run clawbench run --corpus v2 --model your-model to place your agent on the V2 (Hermes) board alongside claude-opus-4-7 and gpt-5.5.
Outcome: You get Intercepted, Reward, Reward (strict), cost per task and Pass/Total (e.g. 58/130) on the same row, plus a downloadable 5-layer trace bundle per run.
An agent passes Intercepted at 54.6% but only reaches 44.6% Reward. You open the trace browser and step through recording.mp4, actions.jsonl, agent-messages.jsonl and requests.jsonl on one clock.
Outcome: You can click a video frame, find the exact action event and LLM turn behind it, and pinpoint whether the agent picked the wrong item, dropped the 'no peanuts' note, or submitted the wrong address.
You download the 918 V1 and 806 V2 frontier-model trajectories and mine success-versus-failure pairs across 13 models on identical tasks.
Outcome: You have JSONL-native data ready for SFT, DPO or PRM without building your own task corpus or alignment pass.
Use Cases
- Evaluate a browser agent on live tasks like booking flights, ordering food, or applying for jobs.
- Compare agent performance with deterministic HTTP interception plus LLM-based reward scoring.
- Submit your model to the public leaderboard and track cost per task against 13 frontier models.
- Download 5-layer trace bundles to diagnose failure modes and audit judge calls.
- Run ClawBench or scope peers with a single CLI command.
- Fine-tune your own agent on 918 V1 + 806 V2 JSONL trajectories with SFT/DPO/PRM.
- Reproduce a leaderboard cell by re-running our judge on your data or your judge on ours.
- Browse all 283 task definitions in a searchable, filterable gallery without downloading.
Models Under the Hood
as of 2026-10-10
Limitations
- ClawBench is a CLI-first benchmark: setup means pip install clawbench-eval and commands like clawbench run --corpus v2 --model your-model.
- The V2 (Hermes) board tracks 130 tasks across 63 live platforms, while the dataset-wide figures are 283 tasks and 163 platforms across all corpora — the differing denominators make eyeball comparisons between V1 and V2 tables easy to misread.
- Stage 2 scoring relies on an LLM judge (deepseek/deepseek-v4-pro) with lenient and strict rubrics, so the headline Reward number carries judgment variability that deterministic interception alone doesn't; claude-opus-4-7 drops from 44.6% Reward to 24.6% under the strict rubric.
- Leaderboard coverage thins out in places: the V2 (OpenClaw) table has a single model at 0.0%, and V2 (Codex) and V2 (Claude Code) are empty at this snapshot.
- Weekly refresh means a pinned number is only current as of its snapshot date (last 2026-05-20, when the corpus stood at 283 tasks and 163 live platforms).
- Because tasks run against live third-party sites, results are volatile by design and old traces won't reproduce byte-for-byte.
as of 2026-10-10
Verification history
We have re-verified ClawBench 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published ClawBench tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0
Ideal for
AI researchers, agent developers and eval teams who need live-site browser-agent numbers they can re-run and audit, not a hosted score.
What this tier adds
Starting tier — the whole benchmark: Apache-2.0 corpus, clawbench-eval CLI, 1,724 runs, 5-layer trace bundles and 918 V1 + 806 V2 JSONL trajectories.
Where the pricing makes sense
The company stage and team size where ClawBench's pricing actually pencils out — and where peers do it cheaper.
ClawBench itself is an Apache-2.0 open-source benchmark with the full corpus, leaderboard and trace bundles published; your real spend is model inference during evaluation runs. That makes it cheaper than commercial eval vendors who charge per seat or per run, but more expensive than a pure sandbox benchmark like WebArena or Mind2Web, where tasks are frozen and no live-site inference or navigation cost is incurred.
Setup time & first value
How long it actually takes to get something useful out of ClawBench — broken out by persona, not the marketing-page minute.
Researchers with Python experience: roughly 15-30 minutes to pip install clawbench-eval, pull a corpus and complete a first run against a model. Agent developers wiring in their own harness: a few hours to a day, since you supply the browser, driver and agent loop around the evaluator. Teams expecting a hosted UI: no path — it's a CLI plus a trace browser.
Switching to or from ClawBench
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From WebArena: port your WebArena-style task definitions onto the ClawBench evaluator, then re-point the same agent loop at live sites and compare Intercepted and Reward against the published V2 (Hermes) board.
- →From Mind2Web: keep your frozen-snapshot pass rates as a baseline and add ClawBench runs on the 130 V2 tasks to see which of your gains survive live-site volatility.
- →From claw-eval / WildClawBench / ClawMark: reuse the existing adapters to score your runs on the ClawBench evaluator without rewriting the harness.
- ↗To WebArena: freeze the ClawBench task list into a sandbox if you need reproducible, byte-identical runs and can accept memorizable tasks.
- ↗To Mind2Web: move to static snapshot evaluation when you need a stable denominator and don't care about live checkout or application flows.
- ↗To WebVoyager-style suites: pick a smaller public suite when 130 tasks and a two-rubric LLM judge is more machinery than your question needs.
Integrations
Resources & Guides
- Resourceclaw-bench.com
Home · ClawBench
Helpful link from claw-bench.com
- Resourceclaw-bench.com
Tasks · ClawBench
Helpful link from claw-bench.com
- Resourceclaw-bench.com
Traces · ClawBench
Helpful link from claw-bench.com
- Resourceclaw-bench.com
Compare · ClawBench
Helpful link from claw-bench.com
- Resourceclaw-bench.com
Contribute · ClawBench
Helpful link from claw-bench.com
Tutorials & Learning
YouTube returned 6 videos for “ClawBench”, and we withheld 6: 6 could not be judged, because “ClawBench” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about ClawBench.
Official links
Tools that pair well with ClawBench
Common stack mates teams adopt alongside ClawBench, with the specific reason each pairing earns its keep.
ClawX
Free, open-source desktop AI assistant that runs 24/7 on your own machine, scraping and analyzing web sources autonomously.
Steel Browser
Open-source cloud browser API for running fleets of AI agent browser sessions at scale
Arena AI
A free public AI leaderboard where live head-to-head votes and guided tasks rank chat models, coding agents, and fullstack builds.
Featured Head-to-Head Comparisons
Clawbench vs Truleo
Truleo and ClawBench serve entirely different worlds. Truleo is a paid, all-in-one intelligence platform for law enforcement, connecting siloed data (jail calls, body cameras, RMS) to generate leads and slash report writing time. ClawBench is a free, open-source benchmark for AI developers to test browser agents on live web tasks. Choose based on your domain: police work or AI research.
Clawbench vs Praktika
Praktika and ClawBench serve entirely different domains. Praktika is a mobile language learning app that uses AI tutors for conversational practice, while ClawBench is an open-source benchmark for evaluating browser-based AI agents on live web tasks. Choose based on your need: improve your spoken English or test an agent's real-world performance.
Clawbench vs Presto Voice
Choose Presto Voice if you run a QSR chain and need a proven drive-thru voice AI to boost revenue and efficiency. Choose ClawBench if you develop or evaluate browser AI agents and need a free, open-source benchmark with real live tasks. These tools serve entirely different purposes—Presto Voice is a production automation platform, ClawBench is a research benchmark.
Alternatives to ClawBench
View allClawX
Free, open-source desktop AI assistant that runs 24/7 on your own machine, scraping and analyzing web sources autonomously.
Steel Browser
Open-source cloud browser API for running fleets of AI agent browser sessions at scale
Frequently Asked Questions
Used ClawBench? Help shape our editorial sentiment research.