Langfuse
Open-source LLM observability, prompt management, and evaluation for teams running AI agents in production.
If data portability and self-hosting matter to you, Langfuse is the most defensible pick in LLM observability: MIT licensed, OpenTelemetry-native, and fast enough after v4 that the old "open source is slower" argument no longer holds. The unified tracing/prompts/evals loop plus an MCP server, CLI 1.0, and SKILL.md for coding agents means one platform covers work you'd otherwise stitch together. Budget carefully though: Cloud starts at $0/mo Hobby (50k units, 30-day retention, 2 users), rises to $29/mo Core, $199/mo Pro, and $2,499/mo Enterprise, and every paid tier includes 100k units then meters at $8/100k. Compare against LangSmith and Braintrust if you want a managed-only experience.
Verified 1d ago · liveness 89/100 · cite: rightaichoice.com/tools/langfuse
- AI engineering teams running multi-turn chat or coding agents in production
- Enterprises needing self-hosting or a US/EU/JP data region for SOC 2, ISO 27001, or HIPAA
- Teams that want tracing, prompts, datasets, experiments, and human review on one data model
- Organizations using IDE coding agents (Claude Code, Cursor) that want prompts and traces via MCP, CLI, or SKILL.md
- Solo developers who want a zero-config logger and won't run ClickHouse, Redis, and blob storage for self-hosting
- Teams expecting flat-rate pricing — Core, Pro, and Enterprise include 100k units/month and then meter at $8/100k units
- Buyers who want a fully managed black box and no interest in data portability or inspecting the source
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Langfuse if you want a flat, unlimited-seat observability bill — every self-serve tier includes only 100k units/month and then meters at $8/100k units, and SSO plus fine-grained RBAC sit behind a $300/mo add-on.
Every paid tier includes only 100k units/month before metering kicks in at $8/100k units — heavy trace volume on verbose agents adds up quickly.
Langfuse Cloud runs $0/mo Hobby (50k units, 30-day retention, 2 users) → $29/mo Core → $199/mo Pro → $2,499/mo Enterprise on a yearly commitment, plus a $300/mo Teams add-on for SSO and RBAC. That fits funded engineering teams with real agent traffic. Hobbyists and pre-revenue prototypes should stay on Hobby or self-host the MIT-licensed core instead of paying for Core.
In short
Langfuse — Open-source LLM observability, prompt management, and evaluation for teams running AI agents in production. Best for AI engineering teams running multi-turn chat or coding agents in production, Enterprises needing self-hosting or a US/EU/JP data region for SOC 2, ISO 27001, or HIPAA, Teams that want tracing, prompts, datasets, experiments, and human review on one data model. Free to start; paid plans from $29/mo.
What's new in Langfuse
Checked yesterdayAcross the latest 5 updates: 4 feature updates and 1 launch.
Run evaluators on historical observations
When you attach an evaluator to a rule, Langfuse now backfills scores onto recent historical observations so production data collected before the evaluator existed still gets graded.
Analyze thousands of observations with the Assistant
The Langfuse Assistant can build a dataset or dashboard, or run code over thousands of observations in a sandbox, with runs continuing in the background while you work elsewhere.
Create alerts for evaluators
Score and cost alerts can now be created directly from evaluator pages, and every alert linked to an evaluator is reviewable from that same page.
Build multi-message prompts and evaluate multi-modal inputs
Two additions to LLM-as-a-Judge evaluators: support for multi-message prompts and the ability to score multi-modal inputs such as images and audio.
Langfuse v4 is live: faster at scale, with more ways to search, monitor, and evaluate
Langfuse v4 ships real-time processing with initial table loads in milliseconds, marketed as up to 165x faster, and at least 10x faster dashboards for large projects.
What people actually say about Langfuse — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
18 mentions across 6 sources (Hacker News, YouTube, Product Hunt, Stack Overflow, GitHub, Lemmy), 55 more we could not attribute · researched Sep 9, 2026.
Weighted by the 73 posts each of 6 sources contributed.
- +Open-source and MIT-licensed, avoiding vendor lock-in with self-hosting.
- +Hierarchical traces give great depth for debugging LangChain and LangGraph.
- +Recent v4 boasts up to 165x faster real-time processing with ClickHouse.
- +LLM-as-a-judge and human annotation queues are easy to set up.
- +100+ integrations cover most frameworks and providers.
- −Learning curve for beginners; UI/UX less intuitive than some alternatives.
- −Initial setup can be complex, especially for self-hosting.
- −Documentation sometimes too high-level; concrete examples missing.
- −Native SDKs don't cover all languages, requiring extra components.
- −Some users still miss error-specific tracking features.
- • If self-hosting, you incur infrastructure costs (ClickHouse, etc.).
- • Custom retention and advanced features may require enterprise plan.
Viability Score
How well maintained and how widely used is Langfuse? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Hierarchical tracing of LLM calls, tool invocations, and retrieval steps
- Session tracking for multi-turn conversations and agentic workflows
- Per-user token and cost tracking for multi-tenant billing
- Agent graphs visualizing complex agentic workflows
- Responsive Trace Timeline with map-style zoom and colour-coded observation types
- LLM-as-a-judge evaluators, including multi-modal and multi-message prompt support
- Code/heuristic evaluators and custom evaluation scores
- Backfill evaluator scores onto historical observations when attaching an evaluator to a rule
- Human annotation queues for building golden datasets
- Evaluator versioning, restore-as-draft, and template starters for chatbots and coding agents
- Prompt versioning, labels, one-click deployments, and rollbacks
- Prompt composability with server- and client-side prompt caching
- Playground for testing prompts on real production inputs and comparing models
- Datasets and Experiments via SDK or UI with side-by-side result comparison
- Langfuse Assistant runs code over thousands of observations in a background sandbox
About Langfuse
Langfuse is an open-source (MIT-licensed) AI engineering platform for tracing, evaluating, and improving LLM and agent applications. Hierarchical traces capture every LLM call, tool invocation, and retrieval step, and you can filter by user, session, cost, latency, or custom metadata. It's OpenTelemetry-native, so it works with any language or framework that supports OTel instrumentation (Python, TypeScript, Go, Java, .NET, Ruby, PHP, Swift), plus 100+ framework and model-provider integrations including LangChain, Vercel AI SDK, LiteLLM, Pydantic AI, Google ADK, CrewAI, LiveKit, OpenAI, Anthropic, Amazon Bedrock, Azure OpenAI, Mistral AI, Google Gemini, xAI, vLLM, and Groq. The platform covers the full engineering loop, not just logging: prompt management with versioning, labels, one-click deploys and rollbacks; a Playground that tests prompts against real production inputs and compares models side by side; datasets and experiments with side-by-side result comparison; and evaluation spanning LLM-as-a-judge, heuristic/code evaluators, user feedback, and human annotation queues. Recent releases push on scale and agent-native workflows. Langfuse v4 is marketed as up to 165x faster, with real-time processing and initial table loads in milliseconds. Attaching an evaluator to a rule now backfills scores onto historical observations. The in-app Langfuse Assistant builds datasets or dashboards and runs code over thousands of observations in a sandbox, with runs continuing in the background. There's a responsive Trace Timeline with map-style zoom, CLI 1.0, a SKILL.md for coding agents, an MCP server that lets IDE agents manage prompts and query traces, and a built-in interactive API reference for self-hosted deployments including air-gapped installs. Underneath is a ClickHouse-backed OLAP architecture with async ingestion and S3/blob storage for large payloads. Deployment is flexible: Langfuse Cloud in a US, EU, or JP region, or self-hosting via Docker Compose, Kubernetes (Helm), and Terraform on AWS, GCP, or Azure, with SOC 2 Type II, ISO 27001, and a HIPAA-ready region. Note that Langfuse joined ClickHouse, the company behind its analytical query layer.
Behind the Verdict
Langfuse's core strength is that it treats observability, prompt management, evaluation, and experiments as one product on one data model rather than four bolted-together tools. A trace you captured can be turned into a dataset, run through an experiment, scored by an LLM-as-a-judge evaluator, reviewed in an annotation queue, and alerted on via Slack or webhook — all without exporting data anywhere. The observability layer itself is detailed: hierarchical traces cover LLM and non-LLM calls including retrieval, embedding and API calls; sessions group multi-turn conversations; agent graphs visualize complex agentic flows; and timelines help you isolate latency problems. Cost and token tracking is per-user, which makes multi-tenant billing tractable. On the evaluation side, evaluators can be attached to rules and now backfill scores onto historical observations, so you don't lose signal from data collected before the evaluator existed; evaluators can be versioned, restored as drafts, created from templates for chatbots, topic detection, exact matches and coding agents, and managed through stable ID-based public APIs. Sampling is consistent across evaluators so comparisons are apples-to-apples. Recent work leans hard into agent-native tooling: the Langfuse Assistant investigates production data and takes approved actions inside the app, runs code over thousands of observations in a background sandbox, and can build datasets and dashboards on request. Coding agents get in through an MCP server, CLI 1.0 (10x+ faster invocations, agent-readable exit codes) and a SKILL.md, so prompts and traces are manageable from an IDE. Multi-modal support extends to images, audio and video in traces, and LLM-as-a-judge evaluators handle multi-modal inputs and multi-message prompts. The operational story is where it diverges from managed competitors. Self-hosting runs on Docker Compose, Kubernetes via Helm, or Terraform on AWS, GCP or Azure, all features are MIT licensed, and self-hosted deployments get a built-in interactive API reference that works even air-gapped. There's an instance switcher for juggling multiple deployments and organized feature previews for admins. The tradeoff is real: if you self-host, you're operating ClickHouse, Redis and blob storage yourself, which is a genuine platform-engineering commitment. Langfuse Cloud removes that, but every paid tier meters: Core, Pro and Enterprise include 100k units/month and then charge $8/100k units, with volume discounts available through the pricing calculator. Data retention is tiered — 30 days on Hobby, 90 days on Core, 3 years on Pro — and SSO (e.g. Okta), SSO enforcement, fine-grained RBAC and a dedicated Slack/MS Teams channel sit in the $300/mo Teams add-on on top of Pro. Audit logs, SCIM provisioning, custom rate limits and SLAs are Enterprise-only at $2,499/mo on a yearly commitment. API rate limits also vary by tier (Core general API 30 requests/min, datasets API 100 requests/min) and the Hobby
Researching Langfuse? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Langfuse actually fits — and what changes day-one when you adopt it.
They instrument the agent with the native Python SDK, capture hierarchical traces with session and userId metadata, and open the Trace Timeline to find the retrieval step adding 4 seconds of latency on failing conversations.
Outcome: They fix the retrieval path, add an LLM-as-judge evaluator on resolution quality through a template, and get alerted in Slack when scores dip.
They deploy Langfuse self-hosted on Kubernetes via Helm in their own cluster, wire traces through OpenTelemetry, and use the built-in interactive API reference to validate the ingestion endpoints.
Outcome: All LLM observability data stays in their infrastructure under the MIT license, with SOC 2 Type II and ISO 27001 reports available when they move to a supported tier.
They connect their IDE agent to the Langfuse MCP server and SKILL.md, then ask the agent to pull failing traces and draft a new prompt version against the linked dataset.
Outcome: Prompt changes are versioned, tested in the Playground and experiments, and deployed via labels without leaving the editor.
Use Cases
- Debug a production agent's unexpected behavior by replaying the exact trace in the Langfuse UI.
- Compare two prompt versions on a dataset of 100 real conversations and pick the winner.
- Wire LLM-as-judge evals into CI to catch quality regressions before shipping a prompt change.
- Track per-user LLM cost in a multi-tenant SaaS and bill accurately.
- Run experiments to compare model providers side by side on your own test cases.
- Build golden datasets via human annotation queues to fine-tune or evaluate models.
- Set up alerts when evaluator scores or cost spike outside expected ranges.
- Backfill judge scores onto historical observations so old production data isn't lost when you add a new evaluator.
Models Under the Hood
as of 2026-09-14
Limitations
- The free Hobby plan includes 50k units/month, 30 days of data access, 2 users and 1 annotation queue, with 1,000 requests/min ingestion throughput — workable for POCs, tight for production.
- Paid self-serve tiers start at $29/mo (Core) and $199/mo (Pro), both including 100k units/month with additional usage at $8/100k units, lower with volume.
- Data retention is tiered: 30 days Hobby, 90 days Core, 3 years Pro.
- SSO, SSO enforcement, and fine-grained RBAC require the $300/mo Teams add-on on top of Pro.
- Audit logs, SCIM provisioning, custom rate limits, and uptime/support SLAs are Enterprise-only at $2,499/mo on a yearly commitment.
- Self-hosting via Docker Compose, Kubernetes (Helm), or Terraform on AWS, GCP or Azure means operating your own infrastructure.
- General API rate limits are 30 requests/min on Core-level plans.
as of 2026-09-29
Verification history
We have re-verified Langfuse 18 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 18 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Langfuse tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Hobby
$0/mo
Ideal for
Solo builders and small teams validating an AI agent idea, running a POC, or instrumenting a side project without a credit card.
What this tier adds
Free entry point: all platform features with limits, 50k units/month, 30-day data access, 2 users, 1 annotation queue, 2 alerts.
Core
$29/mo
Ideal for
A funded team running its first production agent that needs longer history and unlimited collaborators.
What this tier adds
Adds 100k units/month (then $8/100k units), 90-day data access, unlimited users, and in-app support.
Pro
$199/mo
Ideal for
Scaling teams with compliance requirements — SOC 2 and ISO 27001 reports, HIPAA-ready region — and enough trace volume to need higher API ceilings.
What this tier adds
Adds 3-year data retention, data retention management, unlimited annotation queues, high rate limits, and prioritized in-app support.
Teams Add-on
$300/mo
Ideal for
Organizations with an identity provider (Okta, AzureAD/EntraID) that need SSO and role-based access before rolling Langfuse out company-wide.
What this tier adds
Layers SSO, SSO enforcement, fine-grained RBAC, and a dedicated Slack/MS Teams support channel on top of Pro for $300/mo.
Enterprise
$2499/mo
Ideal for
Large regulated organizations needing audit trails, automated provisioning, contractual SLAs, and architecture review support.
What this tier adds
Adds audit logs, SCIM API for user provisioning, custom rate limits, uptime and support SLAs, a named lead support engineer, and yearly commitment with custom volume pricing.
Where the pricing makes sense
The company stage and team size where Langfuse's pricing actually pencils out — and where peers do it cheaper.
Langfuse Cloud runs $0/mo Hobby (50k units, 30-day retention, 2 users) → $29/mo Core → $199/mo Pro → $2,499/mo Enterprise on a yearly commitment, plus a $300/mo Teams add-on for SSO and RBAC. That fits funded engineering teams with real agent traffic. Hobbyists and pre-revenue prototypes should stay on Hobby or self-host the MIT-licensed core instead of paying for Core.
Setup time & first value
How long it actually takes to get something useful out of Langfuse — broken out by persona, not the marketing-page minute.
Sending a first trace from the Python or TypeScript SDK takes minutes: install, set the two API keys, and calls start landing in the UI. A meaningful production setup — OpenTelemetry or framework integration, prompt migration, evaluators, and alerts — is typically a day or two of engineering work. Full self-hosting on Kubernetes with Helm, ClickHouse, Redis, and blob storage is a multi-day
Switching to or from Langfuse
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From LangSmith: re-instrument via OpenTelemetry or the native Python/JS SDKs, then export datasets as JSON and recreate them in Langfuse Experiments.
- →From Helicone or a gateway proxy: keep logging via LiteLLM or the gateway, and add Langfuse as the trace destination rather than rewriting instrumentation.
- →From a homegrown logging table: map existing call records to Langfuse trace/observation structure using the public API or batch import.
- →From spreadsheets of manual eval results: seed Langfuse Datasets and annotation queues to turn them into reusable golden sets.
- ↗To a managed-only competitor: export traces via the batch export to blob storage or query SDK, then replay through the destination's ingestion API.
- ↗To a self-hosted fork: the MIT license lets you take the source and keep operating, but you lose Cloud's SOC 2 / ISO 27001 reports and HIPAA-ready region.
- ↗To a gateway-only logger like LiteLLM proxy logging: drop the SDK, point the gateway at the new sink, and accept that prompt management, datasets and annotation queues have no equivalent.
Integrations
Resources & Guides
- Documentationlangfuse.com
Overview
Langfuse is an open-source LLM engineering platform (GitHub) that helps teams collaboratively debug, analyze, and iterate on their LLM applications. All platform features are natively integrated to accelerate the development workflow.
- Guidelangfuse.com
Guides
End-to-end examples and resources to get started with Langfuse for LLM Tracing, Monitoring, Prompt Management, and more.
- Resourcelangfuse.com
Langfuse Academy
Understand why LLM engineering is different and how to navigate the full AI engineering lifecycle.
- Resourcelangfuse.com
Langfuse
Traces, evals, prompt management and metrics to debug and improve your LLM application.
- Resourcelangfuse.com
Langfuse
Traces, evals, prompt management and metrics to debug and improve your LLM application.
- Resourcelangfuse.com
Support
Overview of available support options for Langfuse.
- Resourcelangfuse.com
Overview
Helpful link from langfuse.com
Tutorials & Learning
YouTube returned 6 videos for “Langfuse”, and we withheld 6: 6 could not be judged, because “Langfuse” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Langfuse.
Tools that pair well with Langfuse
Common stack mates teams adopt alongside Langfuse, with the specific reason each pairing earns its keep.
MLflow
MLflow is the open source AI engineering platform for agent and LLM observability, evaluation, and prompt management.
Phoenix
Open-source tracing, evaluation, and prompt iteration for AI agents — self-host it on your own infrastructure with no per-span bill.
Evidently AI
Open-source AI evaluation and observability for LLMs, RAG, agents, and predictive ML models.
Featured Head-to-Head Comparisons
Langfuse vs Langgraph
Langchain vs Langfuse
If you need deep agent debugging with autonomous failure clustering and fix suggestions, LangSmith is the edge. If you want open-source flexibility, self-hosting, and unified prompt management plus observability, Langfuse is the pragmatic choice. Choose based on whether you need proactive root-cause analysis (LangChain) or full control and compliance via self-hosting (Langfuse).
Langfuse vs Litellm
Langfuse vs Mlflow
Langfuse vs Promptfoo
Alternatives to Langfuse
View allMLflow
MLflow is the open source AI engineering platform for agent and LLM observability, evaluation, and prompt management.
Phoenix
Open-source tracing, evaluation, and prompt iteration for AI agents — self-host it on your own infrastructure with no per-span bill.
Evidently AI
Open-source AI evaluation and observability for LLMs, RAG, agents, and predictive ML models.
Frequently Asked Questions
Used Langfuse? Help shape our editorial sentiment research.