LangWatch Scenario

LangWatch Scenario

Simulation-based AI agent testing that catches failures before production

79/100Safe BetFree planFreemium

LangWatch Scenario is the most integrated testing solution for multi-turn agents with tools or voice. Its combination of an OSS SDK, per-turn judge, and full observability stack is unmatched for production readiness. It's ideal for teams that need automated regression testing beyond single-turn evals. The free Developer tier lets you start immediately. If you only need basic prompt evaluation, consider Langfuse or Arize, but for simulation depth, LangWatch leads.

Verified 4d ago · liveness 79/100 · cite: rightaichoice.com/tools/langwatch-scenario

Best for
  • ML engineers testing multi-turn AI agents with tool calls
  • QA teams automating agent regression in CI/CD
  • Platform teams shipping agent updates with confidence
  • Product managers defining agent behavior via natural language specs
Not ideal for
  • Teams needing only single-turn prompt evaluation
  • Non-technical users who cannot write scenario descriptions
  • Projects requiring closed-source proprietary models only (works with any model but needs API access)
Visit Website

IntermediateFor developers: get a basic scenario running in under 30 minutes using the Python or TypeScript SDK. For PMs: use Langy to generate a test plan in ~14 minutes, then run it in the UI. For voice agents: expect a few hours to configure the RealtimeAgentAdapter and set up integration with your stack.Web · API · CLIAPI availableVerified 4d ago
Pricing
Free plan
FreemiumFree tier3 plans5 hidden costs
Learning curve
Intermediate
For developers: get a basic scenario running in under 30 minutes using the Python or TypeScript SDK. For PMs: use Langy to generate a test plan in ~14 minutes, then run it in the UI. For voice agents: expect a few hours to configure the RealtimeAgentAdapter and set up integration with your stack.
Runs on
WebAPICLI
API available · 15 integrations
Who it's for
ML engineer testing an agent with tool callsQA engineer automating voice agent regressionProduct manager defining agent behavior
Live sentiment
Is LangWatch Scenario actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip LangWatch Scenario if you only need single-turn prompt evaluation, if you want a no-code tool for non-technical users, or if your agent is a simple chatbot without tool use or multi-turn complexity.

The 30-second take
Biggest gripe

Going past the Developer plan's 3 scenarios, 3 simulations, and 3 custom evals requires the Growth plan at €29/core-seat/month.

Price reality

Freemium pricing: Developer at €0/mo is great for solo devs and small teams getting started. Growth at €29/core-seat/month fits growing teams, cheaper than Langfuse Pro (which starts at $50/mo). Enterprise is custom for regulated teams. Compared to Arize or Braintrust, LangWatch's open-source SDK and self-host option can reduce long-term costs if you have the infra.

In short

LangWatch Scenario — Simulation-based AI agent testing that catches failures before production. Best for ML engineers testing multi-turn AI agents with tool calls, QA teams automating agent regression in CI/CD, Platform teams shipping agent updates with confidence. Free to start; paid plans from €29/mo.

What's new in LangWatch Scenario

Checked 2 days ago

Across the latest 10 updates: 5 feature updates, 2 launches and 3 news mentions.

FeatureChangelog·4 days agoNewest

LangWatch 3.17.0: Agent Testing v2 and Prompt Optimization in the Workbench

Agent Testing replaces Simulations; comparison runs arrive; Langy optimizes prompts in the workbench.

FeatureChangelog·11 days ago

LangWatch 3.16.0: New Navigation and Voice APIs support on the Gateway

Sidebar becomes product switcher and icon rail; realtime voice bills to virtual key.

FeatureChangelog·18 days ago

LangWatch 3.13.0: Governed SQL Workbench and Go SDK

Query data in governed SQL workbench; instrument Go apps.

NewsBlog·20 days ago

AI Governance: Can you list every AI agent running in your company right now?

Discusses AI governance and tracking all AI agents in an organization.

NewsBlog·22 days ago

One trace, two layers: OpenTelemetry between your LLM app and your cache

Co-written with BetterDB: how cache decisions and LLM observability layer together.

NewsBlog·22 days ago

The 8 Best LLM Gateways in 2026: Compared for Production

Comparison of LLM gateways for production use; positions LangWatch's gateway.

FeatureChangelog·25 days ago

LangWatch 3.11.0: Coding Agent Capture and Gateway Billing

Capture Copilot and Claude Code sessions; bill gateway spend with webhooks and reconciliation.

FeatureChangelog·Aug 2

LangWatch 3.8.0: The AI Gateway Grows Budgets

Budgets on every dimension of gateway traffic; live spend and redesigned key drawer.

LaunchBlog·Jul 31

Launching Claude Code usage tracking: see where your tokens go

New feature to track Claude Code token usage and spend.

LaunchBlog·Jul 22

Introducing Langy: Your Automated AI Engineer

Langy reads your traces, writes tests, and opens pull requests automatically.

What people actually say about LangWatch Scenario — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

20 mentions across 2 sources (Hacker News, YouTube) · researched Jul 3, 2026.

35% positive65% critical
Recurring strengths
  • +Simulates multi-turn conversations realistically using LLM-powered user simulator.
  • +Each turn judged automatically with pass/fail criteria, surfacing concrete failures.
  • +Open-source SDK (MIT) works with any LLM and any agent framework.
  • +Integrated with LangWatch's full observability stack for traces and metrics.
  • +Supports adversarial red-teaming like Crescendo escalation out of the box.
Recurring frustrations
  • Extremely limited community feedback—only the team's own posts visible.
  • No independent reviews or real-world reliability data yet.
  • YouTube returned zero relevant content; low awareness outside HN.
  • Cost of running judge-agent simulations could add up at scale.
  • Setting up high-quality simulators and judges requires careful prompt work.
Patterns worth knowing
Simulation-based testing is a needed methodology for complex agents.
Seen on Hacker News
The tool is presented as analogous to self-driving car simulation testing.
Seen on Hacker News
Community validation is almost nonexistent—only the creators posting.
Seen on Hacker News
Learning curve
beginnerProductive in ~A few hours
Hidden costs people mention
  • Cost of LLM API calls for user simulator and judge per simulation turn.
  • Cloud platform pricing not publicly visible.
  • Potential data storage costs for trace storage at scale.

Viability Score

79/100
Safe Bet

How well maintained and how widely used is LangWatch Scenario? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
35
What the vendor publishes
60

Last calculated: September 2026

How we score →

Key Features

  • LLM-powered user simulator for realistic multi-turn messages
  • Multi-turn conversation testing with per-turn judge verdicts
  • Configurable success criteria in natural language
  • Tool-call verification across long dialogues
  • Framework-agnostic adapters (LangGraph, CrewAI, Pydantic AI, etc.)
  • Run locally or in CI/CD via pytest/vitest
  • Simulation visualizer with real-time replay
  • Pause, evaluate & annotate mid-conversation
  • Voice agent testing with latency metrics and noise injection
  • Adversarial red-teaming (Crescendo escalation, refusal detection)
  • Open-source Scenario SDK (Python + TypeScript, MIT)
  • Langy: AI assistant generates test plans from plain-English goals
  • Online evaluations and monitors for production traffic
  • Multi-modal evaluations (images and mixed media)
  • Built-in evals (RAGAS, hallucination, toxicity, PII)

About LangWatch Scenario

FreemiumIntermediateAPI availableWeb · API · CLI

LangWatch Scenario is an open-source (MIT) agent testing framework that simulates realistic multi-turn conversations between an LLM-powered user simulator and your AI agent. It evaluates each turn with a judge agent, surfaces failures like wrong tool calls or policy violations, and links every result to a full trace for debugging. The platform caters to developers and PMs building complex AI agents with tool use, voice, or multi-step reasoning. It works with any LLM and any framework (LangGraph, CrewAI, Pydantic AI, etc.) via one-method adapters, and can run locally, in CI/CD, or on LangWatch Cloud. Key components include: an LLM-powered user simulator that generates realistic messages from a scenario description; a judge agent that scores each turn against configurable success criteria; adversarial red-teaming runs (e.g., Crescendo escalation); voice agent simulations with latency metrics and noise injection; and Langy, an AI assistant that turns a PM's goal into a full test plan and drafts prompt revisions. As of April 2026, LangWatch 3.0 offers a simpler self-hosted stack built on ClickHouse (no Elasticsearch required) with Helm or Docker Compose deployment. It also includes a built-in MCP server with OAuth, a Cmd+K command bar for navigation, and the ability to create and run simulations directly in the UI without code. The entire platform is now open source (May 2026). What makes LangWatch Scenario different is its seamless integration with LangWatch's full observability and evaluation stack. Every simulation run is linked to detailed traces, cost/latency metrics, and prompt management. Alternatives like Arize or Langfuse offer observability but lack this depth of simulation-based testing.

Behind the Verdict

LangWatch Scenario earns its place as the go-to choice for teams that live and breathe multi-turn agents. The simulation loop feels like a natural extension of your development process: you write a scenario in plain English, the LLM user simulator pushes your agent through realistic back-and-forths, and the judge agent rules on every turn with a verdict you can trace back to the exact reasoning. That combination of depth and transparency is what sets it apart from observability-only tools like Langfuse or Arize, which will show you what went wrong but won't actively probe your agent with adversarial or voice-based scenarios before your users do.\n\nWhere it really shines is in CI/CD. You can run the same scenarios you craft locally in pytest or vitest and gate merges on the results, so every pull request gets a safety net that catches regressions before they hit production. The Langy assistant takes this further by automating the whole loop: it reads your production traces, drafts a test plan, and even opens a pull request with a prompt fix when the judge flags a problem. In their own dogfooding, they found 14 real bugs this way, including one where the agent got derailed by an unrelated coding question. That's the kind of concrete, proactive testing that manual review just can't match.\n\nBut it's not for everyone. If your product is a simple prompt-response chatbot with no tool calls and no multi-turn complexity, the overhead of scenario simulation is overkill; you'd be better served by a lightweight eval harness. And while the SDK is MIT-licensed, the platform's full power really comes through the cloud or self-hosted LangWatch stack, so you need to be comfortable running infrastructure or paying for the managed tier. Non-technical folks can still define scenarios

Researching LangWatch Scenario? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas LangWatch Scenario actually fits — and what changes day-one when you adopt it.

ML engineer testing an agent with tool calls

You've built a LangGraph agent that books flights. You want to catch a bug where it calls the wrong tool in a long conversation.

Outcome: Write a scenario describing a multi-turn booking flow, run it in pytest or CI, and the judge agent flags the wrong tool call with a full trace for debugging.

QA engineer automating voice agent regression

You have a voice agent using OpenAI Realtime and need to test latency and noise resilience.

Outcome: Set up a voice simulation with the RealtimeAgentAdapter, inject background noise, and get TTFB and p50/p95 latency metrics to ensure quality before launch.

Product manager defining agent behavior

You need to spec out a customer support agent but don't write code.

Outcome: Use Langy to turn your plain-English goal into a test plan, run it in the UI, and review the judge's scores and failures without writing a line of code.

Use Cases

  • Simulate a customer support agent handling complex refund requests over 10 turns
  • Automate regression testing of an AI coding assistant with tool-call verification
  • Red-team a voice agent for security vulnerabilities like prompt injection
  • Generate a test plan from a product manager's plain-English description using Langy
  • Run adversarial simulations to detect policy violations in a financial advisory bot
  • Gate CI/CD releases with simulation-based pass/fail criteria on every commit

Models Under the Hood

Claude CodeCodex

as of 2026-08-28

Limitations

  • The Developer plan limits users to 3 scenarios, 3 simulations, and 3 custom evals, while the Growth plan includes 200k events per month with overage charges of €5 per 100k events.
  • Self-hosted deployment uses ClickHouse and requires Helm or Docker Compose, with no Elasticsearch needed.
  • The platform supports both chat and voice agent testing, and includes red-teaming and governance features, but scaling events and storage beyond included limits incurs additional costs.

as of 2026-08-23

Verification history

We have re-verified LangWatch Scenario 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. re-checked, vendor evidence unchanged
  3. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 7 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Contact sales for a quote
Effective monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published LangWatch Scenario tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Developer

€0/mo

Ideal for

Solo developers or small teams experimenting with agent testing, with up to 3 scenarios and 50k events/month.

What this tier adds

Free starting point: 50k events, 2 users, 3 scenarios, 3 simulations, 3 custom evals, and community support.

Growth

€29 /core-seat/month

Ideal for

Teams shipping agents to production needing unlimited simulations, evals, and prompts, plus support.

What this tier adds

Adds 200k events included (then €5/100k), unlimited lite-users, 30-day retention, and private Slack/Teams support.

Enterprise

Custom

Ideal for

Regulated teams needing on-prem/hybrid deployment, custom SSO/RBAC, audit logs, and SLAs.

What this tier adds

Adds custom retention, ISO 27001 reports, DPA, forward deployed engineer, and billing via AWS/Google Marketplace.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Going past the Developer plan's 3 scenarios, 3 simulations, and 3 custom evals requires the Growth plan at €29/core-seat/month.
  • Beyond the 200k included events on Growth, each additional 100k events costs €5, which adds up at production scale.
  • Extending data retention beyond the included 30 days costs €3 per GB per month, billed only when you exceed 30 days.
  • Enterprise requires a custom contract with volume-based pricing, SSO/RBAC, and dedicated support, which likely means a minimum spend.
  • Self-hosting with ClickHouse and Helm/Docker Compose requires real infrastructure engineering time and resources, which is a cost not captured in the license.

Where the pricing makes sense

The company stage and team size where LangWatch Scenario's pricing actually pencils out — and where peers do it cheaper.

Freemium pricing: Developer at €0/mo is great for solo devs and small teams getting started. Growth at €29/core-seat/month fits growing teams, cheaper than Langfuse Pro (which starts at $50/mo). Enterprise is custom for regulated teams. Compared to Arize or Braintrust, LangWatch's open-source SDK and self-host option can reduce long-term costs if you have the infra.

Setup time & first value

How long it actually takes to get something useful out of LangWatch Scenario — broken out by persona, not the marketing-page minute.

For developers: get a basic scenario running in under 30 minutes using the Python or TypeScript SDK. For PMs: use Langy to generate a test plan in ~14 minutes, then run it in the UI. For voice agents: expect a few hours to configure the RealtimeAgentAdapter and set up integration with your stack.

Switching to or from LangWatch Scenario

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • From Langfuse: import traces via the OpenTelemetry-compatible API and start adding simulations with the Scenario SDK.
  • From Braintrust: use the open-source SDK to run scenarios locally and adopt LangWatch's full observability for production.
Migrating out
  • To Langfuse: if you only need single-turn eval and tracing, Langfuse's API is easy to adopt, but you lose simulation depth.
  • To Arize: for heavy production monitoring, Arize offers a mature observability platform, but you'll need to re-instrument.

Integrations

LangGraphCrewAIPydantic AIClaude CodeElevenLabsOpenAI RealtimeTwilioPipecatGemini LiveOpenTelemetryGitHubSlackTeamsHelmDocker Compose

Resources & Guides

Tutorials & Learning

Tools that pair well with LangWatch Scenario

Common stack mates teams adopt alongside LangWatch Scenario, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to LangWatch Scenario

View all
Future AGI

Future AGI

Simulation-based AI agent testing that catches hallucinations before production

FreemiumTry
CopilotKit

CopilotKit

React frontend stack for building agentic user experiences with generative UI

FreemiumTry
Arena AI

Arena AI

Community-driven leaderboard for comparing AI models, agents, and code through real human votes.

FreemiumTry

Frequently Asked Questions

Used LangWatch Scenario? Help shape our editorial sentiment research.