Deepgram

Deepgram

Real-time speech-to-text, text-to-speech, and voice agent APIs for developers.

95/100Safe BetFree · from $4K+/yearFreemium

Deepgram is a top pick for developers who want a single, low-latency API for STT, TTS, and voice agents. The unified Voice Agent API cuts integration effort, but for simple batch transcription, Nova-3 alone is more cost-effective. If you need an out-of-the-box UI, look elsewhere.

Verified 8d ago · liveness 95/100 · cite: rightaichoice.com/tools/deepgram

Best for
  • Developers building real-time voice agents with the unified Voice Agent API
  • Contact centers needing live transcription and call analytics
  • Global apps requiring multilingual speech recognition (Flux Multilingual)
  • Enterprises that want custom speech models and self-hosting
Not ideal for
  • Teams needing an out-of-the-box UI without coding
  • Low-budget hobbyists who need a perpetual free tier
  • Solo developers who won't benefit from volume discounts
Visit Website

AdvancedDevelopers: under 30 minutes to get a live transcription demo running via the Playground or SDK; a full voice agent with the Voice Agent API typically takes a day to wire up. Batch transcription: minutes to first result. Platform/enterprise: self-hosting or custom models require more time—expect days to weeks depending on infrastructure.APIAPI available6.3k viewsVerified 8d ago
Pricing
Free · from $4K+/year
FreemiumFree tier3 plans6 hidden costs
Learning curve
Advanced
Developers: under 30 minutes to get a live transcription demo running via the Playground or SDK; a full voice agent with the Voice Agent API typically takes a day to wire up. Batch transcription: minutes to first result. Platform/enterprise: self-hosting or custom models require more time—expect days to weeks depending on infrastructure.
Runs on
API
API available · 12 integrations
Who it's for
Developer building a voice agent prototypePlatform engineering lead evaluating speech modelsContact center analytics engineer
Live sentiment
Is Deepgram actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Deepgram if you need a ready-made UI without coding, or if you're a hobbyist relying on a perpetual free tier rather than a one-time $200 credit.

The 30-second take
Biggest gripe

After your $200 free credit runs out, you pay per-minute for STT and per-character for TTS with no monthly cap to protect you from unexpected spikes.

Price reality

Deepgram's pay-as-you-go model fits developers and startups exploring voice AI, with a $200 free credit to start. For high-volume batch transcription, Nova-3 at $0.0048/min undercuts many rivals, though AssemblyAI's similar tier may edge it out on bulk pricing. The Growth tier ($4K+/yr) is best for teams already spending that much monthly—the 20% pre-paid discount is real, but the upfront commitment can sting small teams.

In short

Deepgram — Real-time speech-to-text, text-to-speech, and voice agent APIs for developers. Best for Developers building real-time voice agents with the unified Voice Agent API, Contact centers needing live transcription and call analytics, Global apps requiring multilingual speech recognition (Flux Multilingual). Free to start; paid plans from $4/mo.

What's new in Deepgram

Checked 8 days ago

Across the latest 4 updates: 1 feature update and 3 changelog entries.

What people actually say about Deepgram — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

39 mentions across 4 sources (Hacker News, Product Hunt, Stack Overflow, Lemmy) · researched Aug 18, 2026.

69% positive31% critical
Recurring strengths
  • +Low latency for real-time voice agents (community mentions).
  • +Unified Voice Agent API simplifies STT+TTS+LLM integration.
  • +High accuracy with Nova-3 models, especially multilingual.
  • +Flexible deployment: cloud or self-hosted.
  • +Generous free tier to test and prototype (community acknowledges).
Recurring frustrations
  • Self-hosting setup can be complex and requires resources.
  • Free tier limits may surprise high-volume users.
  • Cloud dependency undermines 'local-first' claims.
  • Documentation could be clearer for beginners (async examples).
  • Advanced features like diarization may be pricey.
Patterns worth knowing
Low latency is a killer feature for real-time voice agents
Seen on Hacker News, Product Hunt, Lemmy
Unified API reduces integration complexity for STT+TTS+LLM
Seen on Hacker News, Lemmy
Multilingual support, especially Flux Multilingual, is a key advantage
Seen on Hacker News, Lemmy
Learning curve
advancedProductive in ~A few hours
Hidden costs people mention
  • Add-ons like speaker diarization may incur extra per-minute charges
  • Self-hosting requires significant infrastructure costs
  • Voice Agent API might charge per call hour, which can add up

Viability Score

95/100
Safe Bet

How well maintained and how widely used is Deepgram? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
69
What the vendor publishes
100

Last calculated: August 2026

How we score →

Key Features

  • Real-time speech-to-text with Flux and Nova-3 models
  • Text-to-speech with Aura-2, Aura-1, and Flux TTS voices
  • Unified Voice Agent API (STT+TTS+LLM orchestration)
  • Flux Multilingual: 10 languages in a single model
  • Batch transcription for pre-recorded audio
  • Self-hosted deployment option
  • Audio Intelligence API for emotion and sentiment analysis
  • Custom model training for edge-case accuracy
  • Speaker diarization
  • Smart Formatting for punctuation and readability
  • Keyterm Prompting for domain-specific jargon
  • Redaction of PII from transcripts
  • Entity Detection
  • Numerals support (e.g., 'three hundred' → '300')
  • Automatic language detection (Nova-3 Multilingual)

About Deepgram

FreemiumAdvancedAPI availableAPI

Deepgram is a Voice AI platform that offers real-time and batch APIs for speech-to-text (STT), text-to-speech (TTS), and voice agents, designed for developers, product teams, and enterprises building conversational AI, contact center analytics, medical transcription, and voice-enabled apps. The platform's centerpiece is a unified Voice Agent API that combines STT, TTS, and LLM orchestration into a single endpoint, reducing integration complexity, latency, and cost—instead of stitched-together components. For transcription, Deepgram offers Flux models (English and Multilingual) tuned for real-time voice agents with built-in turn detection and interruption handling, and Nova-3 models (monolingual and multilingual) for high-accuracy batch and streaming transcription across 45+ languages. Flux Multilingual, launched in July 2026, supports 10 languages in a single model. The TTS side is powered by Aura-2 and Aura-1 voices, delivering natural, low-latency speech for assistants and conversational AI. Developers get flexible deployment—cloud or self-hosted—along with WebSocket/REST APIs, SDKs for multiple languages, and add-ons like Speaker Diarization, Keyterm Prompting, Smart Formatting, and Redaction. Audio Intelligence API adds emotion and sentiment analysis. For platforms and enterprises, Deepgram offers custom models, partner programs, and enterprise solutions. Compared to alternatives like AssemblyAI or Google Cloud Speech-to-Text, Deepgram emphasizes low latency, a single unified API, and a straightforward pricing model with a free tier. Recent updates include new Nova-3 monolingual models and a /llms.txt endpoint for AI agent documentation indexing.

Behind the Verdict

Deepgram is a developer-first voice AI platform that stands out for its unified Voice Agent API, which combines STT, TTS, and LLM orchestration into one call. This is a genuine advantage over stitching together separate services, and it's especially relevant if you're building real-time conversational agents where low latency matters. The Flux models are purpose-built for conversation—they handle turn-taking and interruptions natively, which is exactly the hard part of voice AI. Flux Multilingual (launched July 2026) extends this to 10 languages in one model, so you don't need separate per-language setups. For simpler transcription jobs, Nova-3 is solid and supports 45+ languages, with add-ons like Speaker Diarization, Keyterm Prompting, Smart Formatting, and Redaction. The pricing is straightforward: pay-as-you-go with a $200 free credit, a Growth tier that saves up to 20% on pre-paid credits, and Enterprise for custom models and self-hosting. What's not great: there's no perpetual free tier—you get $200 credit then pay. Concurrency limits on lower tiers (STT up to 50 REST/150 WSS on PAYG, up to 225 WSS on Growth) could bottleneck high-volume use. And it's strictly API-first; if you want a ready-made UI, Deepgram isn't it. For teams that need a full product out of the box, look at vendors like AssemblyAI or Speechmatics which offer more turnkey solutions.

Researching Deepgram? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Deepgram actually fits — and what changes day-one when you adopt it.

Developer building a voice agent prototype

Sign up, grab a free API key, and use the Voice Agent API to tie together Flux STT, an LLM, and Aura-2 TTS.

Outcome: A working voice agent in an afternoon, with natural turn-taking and interruptions handled by the API rather than custom glue code.

Platform engineering lead evaluating speech models

Compare Nova-3 vs Flux on internal test audio, using the Playground to check accuracy on domain-specific jargon.

Outcome: Choose the right model per use case (batch vs real-time) and estimate per-minute costs before committing to a plan.

Contact center analytics engineer

Stream live calls via WebSocket, enabling Speaker Diarization and Redaction, then run Audio Intelligence for sentiment.

Outcome: Live transcripts with speaker labels and PII scrubbed, feeding dashboards for compliance and sentiment analysis in near real-time.

Use Cases

Models Under the Hood

Flux EnglishFlux MultilingualNova-3 MonolingualNova-3 MultilingualAura-2Aura-1Flux TTS

as of 2026-08-14

Limitations

  • Free tier limited to $200 credit; no perpetual free tier.
  • Concurrency limits on lower tiers: STT up to 50 REST, 150 WSS on Pay-as-you-go, up to 225 WSS on Growth.
  • Self-hosted and custom models may require Enterprise plan.
  • API-first design has a learning curve for non-developers.

as of 2026-08-15

Verification history

We have re-verified Deepgram 16 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. re-checked, vendor evidence unchanged
  3. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 16 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Deepgram tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Pay As You Go

$0/mo ($200 free credit)

Ideal for

Solo developer or startup exploring voice AI prototypes with a $200 free credit, needing full API access without a credit card commitment.

What this tier adds

Starting tier with a $200 free credit, all endpoints available, but limited concurrency (STT 50 REST/150 WSS, TTS 45) and no volume discounts.

Growth

$4K+/year

Ideal for

Growing applications with predictable usage that can pre-pay $4K+ annually to get up to 20% savings and higher concurrency (up to 225 WSS).

What this tier adds

Requires $4K+/year pre-payment, boosts concurrency limits (225 WSS STT, 60 TTS/Voice Agent), and offers discounted per-minute rates.

Enterprise

Contact Sales

Ideal for

Large enterprises needing custom models, self-hosted deployment, advanced security/compliance, and tailored SLAs with high volume.

What this tier adds

Adds custom model training, self-hosting, and dedicated support; pricing is contact-sales only.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • After your $200 free credit runs out, you pay per-minute for STT and per-character for TTS with no monthly cap to protect you from unexpected spikes.
  • The Growth plan requires a $4K+ annual pre-payment to get the 20% discount; if your usage is low, you'll pay more upfront than you save.
  • Self-hosted deployment and custom model training are gated behind the Enterprise tier, which is contact-sales only and likely has a custom price tag.
  • Add-ons like Redaction ($0.0020/min), Keyterm Prompting ($0.0013/min), and Entity Detection ($0.0017/min) are NOT included in the base STT rate—they stack on top of the per-minute transcription cost.
  • Flux TTS is free until September 12, 2026, then jumps to $0.0450/1k characters on PAYG—roughly 50% more expensive than Aura-2, so budget accordingly if you build with it.
  • Concurrency is capped on lower tiers (STT up to 150 WSS on PAYG); if you scale past that, you'll need the Growth plan or higher, which means committing to a $4K+ annual spend.

Where the pricing makes sense

The company stage and team size where Deepgram's pricing actually pencils out — and where peers do it cheaper.

Deepgram's pay-as-you-go model fits developers and startups exploring voice AI, with a $200 free credit to start. For high-volume batch transcription, Nova-3 at $0.0048/min undercuts many rivals, though AssemblyAI's similar tier may edge it out on bulk pricing. The Growth tier ($4K+/yr) is best for teams already spending that much monthly—the 20% pre-paid discount is real, but the upfront commitment can sting small teams.

Setup time & first value

How long it actually takes to get something useful out of Deepgram — broken out by persona, not the marketing-page minute.

Developers: under 30 minutes to get a live transcription demo running via the Playground or SDK; a full voice agent with the Voice Agent API typically takes a day to wire up. Batch transcription: minutes to first result. Platform/enterprise: self-hosting or custom models require more time—expect days to weeks depending on infrastructure.

Switching to or from Deepgram

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • From AssemblyAI: 'Deepgram supports reduced latency through the Flux models and a unified Voice Agent API, but you'll need to adapt to its pricing-per-minute model and possible concurrency caps on lower tiers.'
  • From Google Cloud STT: 'Migrating involves rewriting API calls to Deepgram's REST/WebSocket endpoints; the supported audio formats and features like Speaker Diarization are comparable.'
Migrating out
  • To AssemblyAI: 'AssemblyAI offers a similar per-minute model and add-ons; Deepgram's self-hosting and custom models may be harder to replicate.'
  • To Azure Speech: 'Azure provides a broader ecosystem if you're already on Microsoft, but expects you to handle turn-taking and interruption logic yourself—Deepgram's Voice Agent API does it out of the box.'

Integrations

Amazon ConnectTwilioAsteriskPipecatLiveKitGoogle Dialogflow CXGenesysAudioCodesZapierZoomMake.comAWS S3

Resources & Guides

Tutorials & Learning

Tools that pair well with Deepgram

Common stack mates teams adopt alongside Deepgram, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Deepgram

View all
AssemblyAI

AssemblyAI

Production-grade speech-to-text and voice agent APIs for building voice AI.

FreemiumTry
ElevenLabs

ElevenLabs

ElevenLabs: AI voice platform for text-to-speech, voice cloning, dubbing, and agents

FreemiumTry
Fish Audio

Fish Audio

Free expressive text-to-speech and voice cloning API with emotion control

FreemiumTry

Frequently Asked Questions

Used Deepgram? Help shape our editorial sentiment research.