Langfuse

Langfuse

Open-source LLM observability, prompt management, and evaluation for teams running AI agents in production.

89/100Safe BetFree · from $29/moFreemium

If data portability and self-hosting matter to you, Langfuse is the most defensible pick in LLM observability: MIT licensed, OpenTelemetry-native, and fast enough after v4 that the old "open source is slower" argument no longer holds. The unified tracing/prompts/evals loop plus an MCP server, CLI 1.0, and SKILL.md for coding agents means one platform covers work you'd otherwise stitch together. Budget carefully though: Cloud starts at $0/mo Hobby (50k units, 30-day retention, 2 users), rises to $29/mo Core, $199/mo Pro, and $2,499/mo Enterprise, and every paid tier includes 100k units then meters at $8/100k. Compare against LangSmith and Braintrust if you want a managed-only experience.

Verified 1d ago · liveness 89/100 · cite: rightaichoice.com/tools/langfuse

Best for
  • AI engineering teams running multi-turn chat or coding agents in production
  • Enterprises needing self-hosting or a US/EU/JP data region for SOC 2, ISO 27001, or HIPAA
  • Teams that want tracing, prompts, datasets, experiments, and human review on one data model
  • Organizations using IDE coding agents (Claude Code, Cursor) that want prompts and traces via MCP, CLI, or SKILL.md
Not ideal for
  • Solo developers who want a zero-config logger and won't run ClickHouse, Redis, and blob storage for self-hosting
  • Teams expecting flat-rate pricing — Core, Pro, and Enterprise include 100k units/month and then meter at $8/100k units
  • Buyers who want a fully managed black box and no interest in data portability or inspecting the source
Visit Website

IntermediateSending a first trace from the Python or TypeScript SDK takes minutes: install, set the two API keys, and calls start landing in the UI. A meaningful production setup — OpenTelemetry or framework integration, prompt migration, evaluators, and alerts — is typically a day or two of engineering work. Full self-hosting on Kubernetes with Helm, ClickHouse, Redis, and blob storage is a multi-dayWeb · APIAPI available6.5k viewsVerified 1d ago
Pricing
Free · from $29/mo
FreemiumFree tier5 plans6 hidden costs
Learning curve
Intermediate
Sending a first trace from the Python or TypeScript SDK takes minutes: install, set the two API keys, and calls start landing in the UI. A meaningful production setup — OpenTelemetry or framework integration, prompt migration, evaluators, and alerts — is typically a day or two of engineering work. Full self-hosting on Kubernetes with Helm, ClickHouse, Redis, and blob storage is a multi-day
Runs on
WebAPI
API available · 15 integrations
Who it's for
AI engineer shipping a customer-support chat agentPlatform lead at a company with EU data-residency requirementsEngineer using Claude Code and Cursor in a monorepo
Live sentiment
Is Langfuse actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Langfuse if you want a flat, unlimited-seat observability bill — every self-serve tier includes only 100k units/month and then meters at $8/100k units, and SSO plus fine-grained RBAC sit behind a $300/mo add-on.

The 30-second take
Biggest gripe

Every paid tier includes only 100k units/month before metering kicks in at $8/100k units — heavy trace volume on verbose agents adds up quickly.

Price reality

Langfuse Cloud runs $0/mo Hobby (50k units, 30-day retention, 2 users) → $29/mo Core → $199/mo Pro → $2,499/mo Enterprise on a yearly commitment, plus a $300/mo Teams add-on for SSO and RBAC. That fits funded engineering teams with real agent traffic. Hobbyists and pre-revenue prototypes should stay on Hobby or self-host the MIT-licensed core instead of paying for Core.

In short

Langfuse — Open-source LLM observability, prompt management, and evaluation for teams running AI agents in production. Best for AI engineering teams running multi-turn chat or coding agents in production, Enterprises needing self-hosting or a US/EU/JP data region for SOC 2, ISO 27001, or HIPAA, Teams that want tracing, prompts, datasets, experiments, and human review on one data model. Free to start; paid plans from $29/mo.

What's new in Langfuse

Checked yesterday

Across the latest 5 updates: 4 feature updates and 1 launch.

What people actually say about Langfuse — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

18 mentions across 6 sources (Hacker News, YouTube, Product Hunt, Stack Overflow, GitHub, Lemmy), 55 more we could not attribute · researched Sep 9, 2026.

76% positive24% critical

Weighted by the 73 posts each of 6 sources contributed.

Recurring strengths
  • +Open-source and MIT-licensed, avoiding vendor lock-in with self-hosting.
  • +Hierarchical traces give great depth for debugging LangChain and LangGraph.
  • +Recent v4 boasts up to 165x faster real-time processing with ClickHouse.
  • +LLM-as-a-judge and human annotation queues are easy to set up.
  • +100+ integrations cover most frameworks and providers.
Recurring frustrations
  • −Learning curve for beginners; UI/UX less intuitive than some alternatives.
  • −Initial setup can be complex, especially for self-hosting.
  • −Documentation sometimes too high-level; concrete examples missing.
  • −Native SDKs don't cover all languages, requiring extra components.
  • −Some users still miss error-specific tracking features.
Patterns worth knowing
Strong open-source alternative to proprietary tools like LangSmith
Seen on Hacker News, YouTube, Product Hunt, Lemmy
Recent v4 release with dramatic performance improvements
Seen on Hacker News, YouTube, GitHub
Need for simpler UX and better onboarding
Seen on YouTube, Product Hunt, Hacker News
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • If self-hosting, you incur infrastructure costs (ClickHouse, etc.).
  • • Custom retention and advanced features may require enterprise plan.

Viability Score

89/100
Safe Bet

How well maintained and how widely used is Langfuse? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
73
What the vendor publishes
80

Last calculated: September 2026

How we score →

Key Features

  • Hierarchical tracing of LLM calls, tool invocations, and retrieval steps
  • Session tracking for multi-turn conversations and agentic workflows
  • Per-user token and cost tracking for multi-tenant billing
  • Agent graphs visualizing complex agentic workflows
  • Responsive Trace Timeline with map-style zoom and colour-coded observation types
  • LLM-as-a-judge evaluators, including multi-modal and multi-message prompt support
  • Code/heuristic evaluators and custom evaluation scores
  • Backfill evaluator scores onto historical observations when attaching an evaluator to a rule
  • Human annotation queues for building golden datasets
  • Evaluator versioning, restore-as-draft, and template starters for chatbots and coding agents
  • Prompt versioning, labels, one-click deployments, and rollbacks
  • Prompt composability with server- and client-side prompt caching
  • Playground for testing prompts on real production inputs and comparing models
  • Datasets and Experiments via SDK or UI with side-by-side result comparison
  • Langfuse Assistant runs code over thousands of observations in a background sandbox

About Langfuse

FreemiumIntermediateAPI availableWeb · API

Langfuse is an open-source (MIT-licensed) AI engineering platform for tracing, evaluating, and improving LLM and agent applications. Hierarchical traces capture every LLM call, tool invocation, and retrieval step, and you can filter by user, session, cost, latency, or custom metadata. It's OpenTelemetry-native, so it works with any language or framework that supports OTel instrumentation (Python, TypeScript, Go, Java, .NET, Ruby, PHP, Swift), plus 100+ framework and model-provider integrations including LangChain, Vercel AI SDK, LiteLLM, Pydantic AI, Google ADK, CrewAI, LiveKit, OpenAI, Anthropic, Amazon Bedrock, Azure OpenAI, Mistral AI, Google Gemini, xAI, vLLM, and Groq. The platform covers the full engineering loop, not just logging: prompt management with versioning, labels, one-click deploys and rollbacks; a Playground that tests prompts against real production inputs and compares models side by side; datasets and experiments with side-by-side result comparison; and evaluation spanning LLM-as-a-judge, heuristic/code evaluators, user feedback, and human annotation queues. Recent releases push on scale and agent-native workflows. Langfuse v4 is marketed as up to 165x faster, with real-time processing and initial table loads in milliseconds. Attaching an evaluator to a rule now backfills scores onto historical observations. The in-app Langfuse Assistant builds datasets or dashboards and runs code over thousands of observations in a sandbox, with runs continuing in the background. There's a responsive Trace Timeline with map-style zoom, CLI 1.0, a SKILL.md for coding agents, an MCP server that lets IDE agents manage prompts and query traces, and a built-in interactive API reference for self-hosted deployments including air-gapped installs. Underneath is a ClickHouse-backed OLAP architecture with async ingestion and S3/blob storage for large payloads. Deployment is flexible: Langfuse Cloud in a US, EU, or JP region, or self-hosting via Docker Compose, Kubernetes (Helm), and Terraform on AWS, GCP, or Azure, with SOC 2 Type II, ISO 27001, and a HIPAA-ready region. Note that Langfuse joined ClickHouse, the company behind its analytical query layer.

Behind the Verdict

Langfuse's core strength is that it treats observability, prompt management, evaluation, and experiments as one product on one data model rather than four bolted-together tools. A trace you captured can be turned into a dataset, run through an experiment, scored by an LLM-as-a-judge evaluator, reviewed in an annotation queue, and alerted on via Slack or webhook — all without exporting data anywhere. The observability layer itself is detailed: hierarchical traces cover LLM and non-LLM calls including retrieval, embedding and API calls; sessions group multi-turn conversations; agent graphs visualize complex agentic flows; and timelines help you isolate latency problems. Cost and token tracking is per-user, which makes multi-tenant billing tractable. On the evaluation side, evaluators can be attached to rules and now backfill scores onto historical observations, so you don't lose signal from data collected before the evaluator existed; evaluators can be versioned, restored as drafts, created from templates for chatbots, topic detection, exact matches and coding agents, and managed through stable ID-based public APIs. Sampling is consistent across evaluators so comparisons are apples-to-apples. Recent work leans hard into agent-native tooling: the Langfuse Assistant investigates production data and takes approved actions inside the app, runs code over thousands of observations in a background sandbox, and can build datasets and dashboards on request. Coding agents get in through an MCP server, CLI 1.0 (10x+ faster invocations, agent-readable exit codes) and a SKILL.md, so prompts and traces are manageable from an IDE. Multi-modal support extends to images, audio and video in traces, and LLM-as-a-judge evaluators handle multi-modal inputs and multi-message prompts. The operational story is where it diverges from managed competitors. Self-hosting runs on Docker Compose, Kubernetes via Helm, or Terraform on AWS, GCP or Azure, all features are MIT licensed, and self-hosted deployments get a built-in interactive API reference that works even air-gapped. There's an instance switcher for juggling multiple deployments and organized feature previews for admins. The tradeoff is real: if you self-host, you're operating ClickHouse, Redis and blob storage yourself, which is a genuine platform-engineering commitment. Langfuse Cloud removes that, but every paid tier meters: Core, Pro and Enterprise include 100k units/month and then charge $8/100k units, with volume discounts available through the pricing calculator. Data retention is tiered — 30 days on Hobby, 90 days on Core, 3 years on Pro — and SSO (e.g. Okta), SSO enforcement, fine-grained RBAC and a dedicated Slack/MS Teams channel sit in the $300/mo Teams add-on on top of Pro. Audit logs, SCIM provisioning, custom rate limits and SLAs are Enterprise-only at $2,499/mo on a yearly commitment. API rate limits also vary by tier (Core general API 30 requests/min, datasets API 100 requests/min) and the Hobby

Researching Langfuse? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Langfuse actually fits — and what changes day-one when you adopt it.

AI engineer shipping a customer-support chat agent

They instrument the agent with the native Python SDK, capture hierarchical traces with session and userId metadata, and open the Trace Timeline to find the retrieval step adding 4 seconds of latency on failing conversations.

Outcome: They fix the retrieval path, add an LLM-as-judge evaluator on resolution quality through a template, and get alerted in Slack when scores dip.

Platform lead at a company with EU data-residency requirements

They deploy Langfuse self-hosted on Kubernetes via Helm in their own cluster, wire traces through OpenTelemetry, and use the built-in interactive API reference to validate the ingestion endpoints.

Outcome: All LLM observability data stays in their infrastructure under the MIT license, with SOC 2 Type II and ISO 27001 reports available when they move to a supported tier.

Engineer using Claude Code and Cursor in a monorepo

They connect their IDE agent to the Langfuse MCP server and SKILL.md, then ask the agent to pull failing traces and draft a new prompt version against the linked dataset.

Outcome: Prompt changes are versioned, tested in the Playground and experiments, and deployed via labels without leaving the editor.

Use Cases

  • Debug a production agent's unexpected behavior by replaying the exact trace in the Langfuse UI.
  • Compare two prompt versions on a dataset of 100 real conversations and pick the winner.
  • Wire LLM-as-judge evals into CI to catch quality regressions before shipping a prompt change.
  • Track per-user LLM cost in a multi-tenant SaaS and bill accurately.
  • Run experiments to compare model providers side by side on your own test cases.
  • Build golden datasets via human annotation queues to fine-tune or evaluate models.
  • Set up alerts when evaluator scores or cost spike outside expected ranges.
  • Backfill judge scores onto historical observations so old production data isn't lost when you add a new evaluator.

Models Under the Hood

OpenAIAnthropicAmazon BedrockAzure OpenAIMistral AIGeminixAIvLLM

as of 2026-09-14

Limitations

  • The free Hobby plan includes 50k units/month, 30 days of data access, 2 users and 1 annotation queue, with 1,000 requests/min ingestion throughput — workable for POCs, tight for production.
  • Paid self-serve tiers start at $29/mo (Core) and $199/mo (Pro), both including 100k units/month with additional usage at $8/100k units, lower with volume.
  • Data retention is tiered: 30 days Hobby, 90 days Core, 3 years Pro.
  • SSO, SSO enforcement, and fine-grained RBAC require the $300/mo Teams add-on on top of Pro.
  • Audit logs, SCIM provisioning, custom rate limits, and uptime/support SLAs are Enterprise-only at $2,499/mo on a yearly commitment.
  • Self-hosting via Docker Compose, Kubernetes (Helm), or Terraform on AWS, GCP or Azure means operating your own infrastructure.
  • General API rate limits are 30 requests/min on Core-level plans.

as of 2026-09-29

Verification history

We have re-verified Langfuse 18 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 18 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Langfuse tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Hobby

$0/mo

Ideal for

Solo builders and small teams validating an AI agent idea, running a POC, or instrumenting a side project without a credit card.

What this tier adds

Free entry point: all platform features with limits, 50k units/month, 30-day data access, 2 users, 1 annotation queue, 2 alerts.

Core

$29/mo

Ideal for

A funded team running its first production agent that needs longer history and unlimited collaborators.

What this tier adds

Adds 100k units/month (then $8/100k units), 90-day data access, unlimited users, and in-app support.

Pro

$199/mo

Ideal for

Scaling teams with compliance requirements — SOC 2 and ISO 27001 reports, HIPAA-ready region — and enough trace volume to need higher API ceilings.

What this tier adds

Adds 3-year data retention, data retention management, unlimited annotation queues, high rate limits, and prioritized in-app support.

Teams Add-on

$300/mo

Ideal for

Organizations with an identity provider (Okta, AzureAD/EntraID) that need SSO and role-based access before rolling Langfuse out company-wide.

What this tier adds

Layers SSO, SSO enforcement, fine-grained RBAC, and a dedicated Slack/MS Teams support channel on top of Pro for $300/mo.

Enterprise

$2499/mo

Ideal for

Large regulated organizations needing audit trails, automated provisioning, contractual SLAs, and architecture review support.

What this tier adds

Adds audit logs, SCIM API for user provisioning, custom rate limits, uptime and support SLAs, a named lead support engineer, and yearly commitment with custom volume pricing.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Every paid tier includes only 100k units/month before metering kicks in at $8/100k units — heavy trace volume on verbose agents adds up quickly.
  • SSO via Okta, SSO enforcement, and fine-grained RBAC are not in Pro; they require the $300/mo Teams add-on stacked on top of the $199/mo Pro plan.
  • Audit logs, SCIM user provisioning, custom rate limits, and uptime/support SLAs are locked to Enterprise at $2,499/mo, which carries a yearly commitment.
  • Staying on the free Hobby plan caps you at 30 days of data access and 2 users, so long-horizon trend analysis or team review forces an upgrade.
  • Data retention is tiered (30 days Hobby / 90 days Core / 3 years Pro), so compliance-driven retention requirements push you up the ladder.
  • Self-hosting avoids license fees but leaves you paying for and operating ClickHouse, Redis, and blob storage infrastructure yourself.

Where the pricing makes sense

The company stage and team size where Langfuse's pricing actually pencils out — and where peers do it cheaper.

Langfuse Cloud runs $0/mo Hobby (50k units, 30-day retention, 2 users) → $29/mo Core → $199/mo Pro → $2,499/mo Enterprise on a yearly commitment, plus a $300/mo Teams add-on for SSO and RBAC. That fits funded engineering teams with real agent traffic. Hobbyists and pre-revenue prototypes should stay on Hobby or self-host the MIT-licensed core instead of paying for Core.

Setup time & first value

How long it actually takes to get something useful out of Langfuse — broken out by persona, not the marketing-page minute.

Sending a first trace from the Python or TypeScript SDK takes minutes: install, set the two API keys, and calls start landing in the UI. A meaningful production setup — OpenTelemetry or framework integration, prompt migration, evaluators, and alerts — is typically a day or two of engineering work. Full self-hosting on Kubernetes with Helm, ClickHouse, Redis, and blob storage is a multi-day

Switching to or from Langfuse

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From LangSmith: re-instrument via OpenTelemetry or the native Python/JS SDKs, then export datasets as JSON and recreate them in Langfuse Experiments.
  • →From Helicone or a gateway proxy: keep logging via LiteLLM or the gateway, and add Langfuse as the trace destination rather than rewriting instrumentation.
  • →From a homegrown logging table: map existing call records to Langfuse trace/observation structure using the public API or batch import.
  • →From spreadsheets of manual eval results: seed Langfuse Datasets and annotation queues to turn them into reusable golden sets.
Migrating out
  • ↗To a managed-only competitor: export traces via the batch export to blob storage or query SDK, then replay through the destination's ingestion API.
  • ↗To a self-hosted fork: the MIT license lets you take the source and keep operating, but you lose Cloud's SOC 2 / ISO 27001 reports and HIPAA-ready region.
  • ↗To a gateway-only logger like LiteLLM proxy logging: drop the SDK, point the gateway at the new sink, and accept that prompt management, datasets and annotation queues have no equivalent.

Integrations

LangChainVercel AI SDKLiteLLMPydantic AIGoogle ADKCrewAILiveKitOpenAIAnthropicAmazon BedrockAzure OpenAIMistral AIGoogle GeminixAIPostHog

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Langfuse”, and we withheld 6: 6 could not be judged, because “Langfuse” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Langfuse.

Frequently Asked Questions

Used Langfuse? Help shape our editorial sentiment research.