CUREBench

CUREBench

Open-source benchmark scoring AI reasoning on therapeutic decision-making, with a 2026 software-engineering edition.

62/100MonitorFreeFree

Pick CURE-Bench if you need a defensible, trace-based way to compare reasoning models on therapeutic decisions — the dual-track setup, robustness metrics, and human expert review go well past one-number leaderboards. Skip it if you're shopping for a clinical product or a leaderboard that maps cleanly to deployment readiness; it grades reasoning, not bedside performance. The 2026 software-engineering expansion and CueBench for Developers make it relevant to teams who never touch healthcare.

Verified 12h ago · liveness 62/100 · cite: rightaichoice.com/tools/curebench

Best for
  • AI researchers comparing reasoning models on multi-step therapeutic decisions
  • Healthcare regulators needing trace-level evidence of model reasoning quality
  • Pharma and clinical teams evaluating drug safety, dosing, and repurposing reasoning
  • Academic labs studying clinical AI with reproducible, expert-reviewed metrics
Not ideal for
  • Clinicians who need real-time clinical decision support at the bedside
  • Teams wanting a production-ready therapeutic AI system rather than a benchmark
  • Groups that cannot produce JSONL traces, token usage, and model metadata
Visit Website

AdvancedExpect days to low weeks, not minutes. Track 1 needs the harness plus a submission bundle carrying reasoning traces, token usage, and model metadata. Track 2 adds ToolUniverse provisioning and tool-call logging. CueBench for Developers and the SWE evaluation sit on separate surfaces, so budget extra time if you want all three.Web · CLINo public APIVerified 12h ago
Pricing
Free
FreeFree tier4 hidden costs
Learning curve
Advanced
Expect days to low weeks, not minutes. Track 1 needs the harness plus a submission bundle carrying reasoning traces, token usage, and model metadata. Track 2 adds ToolUniverse provisioning and tool-call logging. CueBench for Developers and the SWE evaluation sit on separate surfaces, so budget extra time if you want all three.
Runs on
WebCLI
No public API · 5 integrations
Who it's for
Academic lab comparing reasoning modelsPharma team evaluating an agentic biomedical pipelineEngineering team measuring human-agent collaboration
Live sentiment
Is CUREBench actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip CURE-Bench if you need a deployed therapeutic AI system or a benchmark that maps directly to bedside deployment readiness, since it grades reasoning traces rather than clinical outcomes.

The 30-second take
Biggest gripe

Submissions require reasoning traces, token usage, and model metadata, so you pay the compute cost of generating and storing traces for every model you enter.

Price reality

CURE-Bench is published as open-source benchmark infrastructure, so the cost question is the compute and engineering time you spend running evaluations, not a subscription. The heavier line item is Track 2: tool-augmented runs against FDA, OpenTargets, and PubMed plus the tool-call logs mean real API and infrastructure spend per model you enter. Budget for that rather than for a licence.

In short

CUREBench — Open-source benchmark scoring AI reasoning on therapeutic decision-making, with a 2026 software-engineering edition. Best for AI researchers comparing reasoning models on multi-step therapeutic decisions, Healthcare regulators needing trace-level evidence of model reasoning quality, Pharma and clinical teams evaluating drug safety, dosing, and repurposing reasoning. Free to use.

What people actually say about CUREBench — is it worth it?

We scanned public community sources for CUREBench on Aug 4, 2026 and could not establish that the discussion we found is about this tool rather than something else sharing its name. Our own analysis of that scan says the posts were off-subject. Rather than publish a sentiment score built on the wrong subject, we publish nothing here and re-run the scan.

Viability Score

62/100
Monitor

How well maintained and how widely used is CUREBench? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
90
Site health
95
User sentiment
10
What the vendor publishes
20

Last calculated: October 2026

How we score →

Key Features

  • Open-source benchmark for AI therapeutic decision-making
  • Track 1: internal model reasoning with no external tools, APIs, or retrieval
  • Track 2: agentic tool-augmented reasoning with FDA, OpenTargets, and PubMed
  • ToolUniverse toolbox provided for agentic tool orchestration
  • 12 real-world biomedical reasoning tasks spanning drug labeling, safety, and regulation
  • Treatment recommendation task considering patient populations
  • Adverse event prediction from drug properties and patient factors
  • Dosage and administration task covering forms, strengths, and instructions
  • Drug warnings and safety task covering boxed warnings, contraindications, and interactions
  • Drug use in specific populations: pregnancy, pediatric, geriatric, nursing mothers
  • Pharmacology task: mechanism of action, pharmacodynamics, pharmacokinetics
  • Nonclinical toxicology task: carcinogenesis, mutagenesis, fertility impairment
  • Patient-focused information task: medication guides and package inserts
  • Agentic dataset pipeline: QuestionGen, TraceGen, and ToolGen
  • Submissions require reasoning traces, final answers, token usage, and model metadata

About CUREBench

FreeAdvancedNo APIWeb · CLI

CURE-Bench is a competition-grade, open-source benchmark for AI reasoning in therapeutic decision-making. It was built for the NeurIPS 2025 workshop in San Diego (Saturday, December 6, 2025, 2:00–4:45 p.m. PST, Upper Level Ballroom 6DE) and organized with Harvard Medical School, MIT, the Kempner Institute, Brigham and Women's Hospital, the Chan Zuckerberg Initiative, and the Milken Institute. It exists because most medical evaluations are QA-style; CURE-Bench instead tests models on the messy, multi-step work of recommending treatments, assessing drug safety and efficacy, designing regimens, and spotting repurposing opportunities. Two competition tracks separate what you're actually measuring. Track 1 scores internal model reasoning with no tools, APIs, or retrieval — the site names LLaMA and DeepSeek-R1 as the kind of standalone models this fits. Track 2 scores agentic reasoning where models orchestrate biomedical tools such as FDA, OpenTargets, and PubMed through the provided ToolUniverse toolbox, with multi-agent pipelines encouraged. Across 12 therapeutic reasoning tasks — treatment recommendation, adverse events, drug warnings and safety, dependence and abuse, dosage and administration, use in specific populations, pharmacology, clinical information, nonclinical toxicology, patient-focused information, drug overview, and drug ingredients — submissions must ship reasoning traces, final answers, token usage, and model metadata, plus a tool-call log in Track 2. Evaluation is a weighted aggregate rather than a single accuracy number: direct performance, accumulated step correctness, open-ended accuracy, rephrasing consistency, option-order robustness, token efficiency, and tool-usage quality. An agentic judge combines factuality checks via retrieval-augmented generation with clinical-relevance scoring, and the top 5–10 teams get a human expert review from clinicians, clinical researchers, and pharmacists — a validity check aimed at catching metric hacking. The dataset pipeline (QuestionGen, TraceGen, ToolGen) generates thousands of clinically grounded examples on top of TxAgent's generation system. In 2026 the project widened beyond clinical work: CUREBench now evaluates 13 models and 4 agents on software-engineering tasks across Go, Java, Python, Rust, and TypeScript, and the CueBench for Developers edition scores how effectively humans drive coding agents. This is research infrastructure, not a clinical product — it grades reasoning, and it does not sit at the bedside.

Behind the Verdict

Most medical AI benchmarks collapse a hard question into one accuracy percentage. CURE-Bench refuses that shortcut, and that refusal is its main asset. If you submit to Track 2, you hand over a reasoning trace, a tool-call log, token usage, and model metadata — so a reviewer can see not just that a model landed on the right drug but whether it got there through a defensible chain of inference or through a lucky guess. The robustness metrics carry real weight here: rephrasing consistency and option-order robustness are exactly the tests that catch a model that memorized answer patterns rather than reasoning about pharmacology. The fact that the top 5–10 teams get human review from clinicians, clinical researchers, and pharmacists closes the loop that pure metric aggregation leaves open. Strengths: the task set is genuinely specific — 12 tasks spanning drug labeling, adverse events, dosage and administration, warnings and safety, pharmacology, nonclinical toxicology, and patient-focused information, drawn from real labeling and regulatory structures. The dual-track split is the right abstraction, because 'does this model know medicine' and 'can this model orchestrate biomedical tools' are different questions with different answers. ToolUniverse lowers the cost of entering Track 2. And the pipeline (QuestionGen, TraceGen, ToolGen) is documented enough that you can understand where the example data came from rather than taking a black-box eval set on faith. Weaknesses: this is infrastructure you run, not a service you call — the seed notes there is no API, so evaluation work means standing up the harness yourself, and that takes technical depth. The domain is narrow by design: therapeutic decision-making, plus the 2026 SWE expansion. Niche clinical areas may be thinly covered because community contributions are still growing. And the evaluation surface is split across CUREBench and swe-rebench.com, with CueBench for Developers as a separate edition, so a team tracking everything may end up integrating across sites. Where it fits: academic labs and AI research teams that need reproducible, expert-reviewed evidence about a reasoning model's clinical judgment, and pharma or regulatory groups that need trace-level evidence rather than a leaderboard rank. In 2026 there is a second audience entirely — engineering teams who want to measure how well a human drives a coding agent, which is a very different question from how well the agent codes alone. Where it doesn't fit: anyone who needs real-time clinical decision support, or who wants a deployed therapeutic system instead of a score.

Researching CUREBench? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas CUREBench actually fits — and what changes day-one when you adopt it.

Academic lab comparing reasoning models

You want to know whether a newly released reasoning model actually handles pharmacology or just pattern-matches on drug names. You enter it into Track 1 with reasoning traces and token usage, then compare its rephrasing-consistency and option-order-robustness scores against the baseline models already on the leaderboard.

Outcome: A weighted aggregate across direct performance, accumulated step correctness, and robustness metrics that shows whether the gain is real reasoning or memorized surface form.

Pharma team evaluating an agentic biomedical pipeline

Your multi-agent pipeline orchestrates FDA labeling, OpenTargets, and PubMed lookups. You submit to Track 2 with a reasoning trace and a tool-call log through ToolUniverse, so judges can score integration quality and tool-usage effectiveness separately from final-answer accuracy.

Outcome: Tool-usage quality scoring that tells you whether the agent is using the right database for the right question, not just whether it landed on the right answer.

Engineering team measuring human-agent collaboration

Your developers use coding agents daily, and you want to know whether they are driving them well. You use the CueBench for Developers edition, which scores how effectively humans drive coding agents across the Go, Java, Python, Rust, and TypeScript tasks in the 2026 SWE expansion.

Outcome: A separate score for the human half of the loop, which a pure agent-capability leaderboard cannot give you.

Use Cases

  • Evaluate AI model performance on treatment selection for complex diseases.
  • Assess AI reasoning with drug-drug interaction detection tasks.
  • Train and benchmark models using standardized clinical scenarios.
  • Compare different AI architectures on therapeutic decision benchmarks.
  • Validate AI adherence to clinical guidelines in simulated cases.
  • Measure how well developers drive coding agents with CueBench for Developers.
  • Test standalone reasoning models in Track 1 without external tool access.
  • Benchmark biomedical tool orchestration quality in Track 2.

Limitations

  • No API access; currently only a benchmark framework.
  • Requires technical expertise to set up and run evaluations.
  • Limited to therapeutic decision-making tasks (plus the 2026 SWE expansion); not a general AI benchmark.
  • The SWE evaluation is hosted separately at swe-rebench.com, and CueBench for Developers is a separate edition, so you may need to integrate across sites.
  • Community contributions are still growing, so coverage in niche clinical areas may be thin.

as of 2026-10-08

Verification history

We have re-verified CUREBench 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-checked, vendor evidence unchanged
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-checked, vendor evidence unchanged
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 8 verification passes.

Free to cite with attribution — this page re-verifies continuously.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Submissions require reasoning traces, token usage, and model metadata, so you pay the compute cost of generating and storing traces for every model you enter.
  • Track 2 evaluation needs a tool-call log, which means provisioning and running ToolUniverse against FDA, OpenTargets, and PubMed — extra infrastructure and API calls beyond the model inference itself.
  • The SWE evaluation lives separately at swe-rebench.com and CueBench for Developers is its own edition, so a team tracking everything may end up building glue code across three surfaces.
  • Standing up the evaluation harness takes engineering time; this is not a hosted scoring service you point at an endpoint.

Where the pricing makes sense

The company stage and team size where CUREBench's pricing actually pencils out — and where peers do it cheaper.

CURE-Bench is published as open-source benchmark infrastructure, so the cost question is the compute and engineering time you spend running evaluations, not a subscription. The heavier line item is Track 2: tool-augmented runs against FDA, OpenTargets, and PubMed plus the tool-call logs mean real API and infrastructure spend per model you enter. Budget for that rather than for a licence.

Setup time & first value

How long it actually takes to get something useful out of CUREBench — broken out by persona, not the marketing-page minute.

Expect days to low weeks, not minutes. Track 1 needs the harness plus a submission bundle carrying reasoning traces, token usage, and model metadata. Track 2 adds ToolUniverse provisioning and tool-call logging. CueBench for Developers and the SWE evaluation sit on separate surfaces, so budget extra time if you want all three.

Switching to or from CUREBench

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a QA-style medical benchmark: replace the single-accuracy leaderboard with CURE-Bench's dual-track setup, which separates internal reasoning from tool-augmented agentic reasoning.
  • →From an ad hoc internal clinical eval: rebuild your prompts against the 12 published therapeutic reasoning tasks so your numbers are comparable to other submissions.
  • →From in-house trace logging: align your reasoning-trace and tool-call-log format with CURE-Bench submission requirements, including token usage and model metadata.
Migrating out
  • ↗To swe-rebench.com: route software-engineering evaluations there, since the SWE track is hosted separately from the clinical CURE-Bench work.
  • ↗To CueBench for Developers: for questions about how humans drive coding agents, use that edition rather than the model-capability tracks.
  • ↗To a general-purpose LLM benchmark: if therapeutic or clinical reasoning is not your focus, CURE-Bench's task set will not map to what you are measuring.

Integrations

KaggleGitHubFDAOpenTargetsPubMed

Resources & Guides

Tutorials & Learning

YouTube returned 3 videos for “CUREBench”, and we withheld 3: 3 could not be judged, because “CUREBench” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about CUREBench.

Official links

Tools that pair well with CUREBench

Common stack mates teams adopt alongside CUREBench, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to CUREBench

View all
Owkin

Owkin

Owkin's K Pro is an autonomous AI scientist agent that informs biopharma R&D, clinical trial and portfolio decisions.

Contact SalesTry
Insilico Medicine

Insilico Medicine

Generative AI drug discovery suite covering target ID, molecule design, biologics engineering and clinical trial prediction.

Contact SalesTry
Cradle Bio

Cradle Bio

AI-guided protein engineering that co-optimizes binding, stability, activity, and expression using your own wet-lab data.

Contact SalesTry

Frequently Asked Questions

Used CUREBench? Help shape our editorial sentiment research.