Deepgram
Real-time speech-to-text, text-to-speech, and voice agent APIs for developers.
Deepgram is a top pick for developers who want a single, low-latency API for STT, TTS, and voice agents. The unified Voice Agent API cuts integration effort, but for simple batch transcription, Nova-3 alone is more cost-effective. If you need an out-of-the-box UI, look elsewhere.
Verified 8d ago · liveness 95/100 · cite: rightaichoice.com/tools/deepgram
- Developers building real-time voice agents with the unified Voice Agent API
- Contact centers needing live transcription and call analytics
- Global apps requiring multilingual speech recognition (Flux Multilingual)
- Enterprises that want custom speech models and self-hosting
- Teams needing an out-of-the-box UI without coding
- Low-budget hobbyists who need a perpetual free tier
- Solo developers who won't benefit from volume discounts
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Deepgram if you need a ready-made UI without coding, or if you're a hobbyist relying on a perpetual free tier rather than a one-time $200 credit.
After your $200 free credit runs out, you pay per-minute for STT and per-character for TTS with no monthly cap to protect you from unexpected spikes.
Deepgram's pay-as-you-go model fits developers and startups exploring voice AI, with a $200 free credit to start. For high-volume batch transcription, Nova-3 at $0.0048/min undercuts many rivals, though AssemblyAI's similar tier may edge it out on bulk pricing. The Growth tier ($4K+/yr) is best for teams already spending that much monthly—the 20% pre-paid discount is real, but the upfront commitment can sting small teams.
In short
Deepgram — Real-time speech-to-text, text-to-speech, and voice agent APIs for developers. Best for Developers building real-time voice agents with the unified Voice Agent API, Contact centers needing live transcription and call analytics, Global apps requiring multilingual speech recognition (Flux Multilingual). Free to start; paid plans from $4/mo.
What's new in Deepgram
Checked 8 days agoAcross the latest 4 updates: 1 feature update and 3 changelog entries.
Numerals Support Now Available for 4 New Languages: Bulgarian, Chinese (Cantonese, Traditional), Malay, and Korean (Monolingual Models)
Expanded Numerals feature to convert spoken numbers to digits in these languages.
Deepgram Self-Hosted August 2026 Release (260812)
New container images for self-hosted deployments.
Introducing Flux Multilingual: One Conversational Speech Model for Global Voice Agents
Flux Multilingual supports 10 languages in a single model with monolingual-grade accuracy.
Changelog update Aug 10, 2026
General changelog update; specific details not captured.
What people actually say about Deepgram — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
39 mentions across 4 sources (Hacker News, Product Hunt, Stack Overflow, Lemmy) · researched Aug 18, 2026.
- +Low latency for real-time voice agents (community mentions).
- +Unified Voice Agent API simplifies STT+TTS+LLM integration.
- +High accuracy with Nova-3 models, especially multilingual.
- +Flexible deployment: cloud or self-hosted.
- +Generous free tier to test and prototype (community acknowledges).
- −Self-hosting setup can be complex and requires resources.
- −Free tier limits may surprise high-volume users.
- −Cloud dependency undermines 'local-first' claims.
- −Documentation could be clearer for beginners (async examples).
- −Advanced features like diarization may be pricey.
- • Add-ons like speaker diarization may incur extra per-minute charges
- • Self-hosting requires significant infrastructure costs
- • Voice Agent API might charge per call hour, which can add up
Viability Score
How well maintained and how widely used is Deepgram? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: August 2026
How we score →Key Features
- Real-time speech-to-text with Flux and Nova-3 models
- Text-to-speech with Aura-2, Aura-1, and Flux TTS voices
- Unified Voice Agent API (STT+TTS+LLM orchestration)
- Flux Multilingual: 10 languages in a single model
- Batch transcription for pre-recorded audio
- Self-hosted deployment option
- Audio Intelligence API for emotion and sentiment analysis
- Custom model training for edge-case accuracy
- Speaker diarization
- Smart Formatting for punctuation and readability
- Keyterm Prompting for domain-specific jargon
- Redaction of PII from transcripts
- Entity Detection
- Numerals support (e.g., 'three hundred' → '300')
- Automatic language detection (Nova-3 Multilingual)
About Deepgram
Deepgram is a Voice AI platform that offers real-time and batch APIs for speech-to-text (STT), text-to-speech (TTS), and voice agents, designed for developers, product teams, and enterprises building conversational AI, contact center analytics, medical transcription, and voice-enabled apps. The platform's centerpiece is a unified Voice Agent API that combines STT, TTS, and LLM orchestration into a single endpoint, reducing integration complexity, latency, and cost—instead of stitched-together components. For transcription, Deepgram offers Flux models (English and Multilingual) tuned for real-time voice agents with built-in turn detection and interruption handling, and Nova-3 models (monolingual and multilingual) for high-accuracy batch and streaming transcription across 45+ languages. Flux Multilingual, launched in July 2026, supports 10 languages in a single model. The TTS side is powered by Aura-2 and Aura-1 voices, delivering natural, low-latency speech for assistants and conversational AI. Developers get flexible deployment—cloud or self-hosted—along with WebSocket/REST APIs, SDKs for multiple languages, and add-ons like Speaker Diarization, Keyterm Prompting, Smart Formatting, and Redaction. Audio Intelligence API adds emotion and sentiment analysis. For platforms and enterprises, Deepgram offers custom models, partner programs, and enterprise solutions. Compared to alternatives like AssemblyAI or Google Cloud Speech-to-Text, Deepgram emphasizes low latency, a single unified API, and a straightforward pricing model with a free tier. Recent updates include new Nova-3 monolingual models and a /llms.txt endpoint for AI agent documentation indexing.
Behind the Verdict
Deepgram is a developer-first voice AI platform that stands out for its unified Voice Agent API, which combines STT, TTS, and LLM orchestration into one call. This is a genuine advantage over stitching together separate services, and it's especially relevant if you're building real-time conversational agents where low latency matters. The Flux models are purpose-built for conversation—they handle turn-taking and interruptions natively, which is exactly the hard part of voice AI. Flux Multilingual (launched July 2026) extends this to 10 languages in one model, so you don't need separate per-language setups. For simpler transcription jobs, Nova-3 is solid and supports 45+ languages, with add-ons like Speaker Diarization, Keyterm Prompting, Smart Formatting, and Redaction. The pricing is straightforward: pay-as-you-go with a $200 free credit, a Growth tier that saves up to 20% on pre-paid credits, and Enterprise for custom models and self-hosting. What's not great: there's no perpetual free tier—you get $200 credit then pay. Concurrency limits on lower tiers (STT up to 50 REST/150 WSS on PAYG, up to 225 WSS on Growth) could bottleneck high-volume use. And it's strictly API-first; if you want a ready-made UI, Deepgram isn't it. For teams that need a full product out of the box, look at vendors like AssemblyAI or Speechmatics which offer more turnkey solutions.
Researching Deepgram? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Deepgram actually fits — and what changes day-one when you adopt it.
Sign up, grab a free API key, and use the Voice Agent API to tie together Flux STT, an LLM, and Aura-2 TTS.
Outcome: A working voice agent in an afternoon, with natural turn-taking and interruptions handled by the API rather than custom glue code.
Compare Nova-3 vs Flux on internal test audio, using the Playground to check accuracy on domain-specific jargon.
Outcome: Choose the right model per use case (batch vs real-time) and estimate per-minute costs before committing to a plan.
Stream live calls via WebSocket, enabling Speaker Diarization and Redaction, then run Audio Intelligence for sentiment.
Outcome: Live transcripts with speaker labels and PII scrubbed, feeding dashboards for compliance and sentiment analysis in near real-time.
Use Cases
- Build real-time voice agents for customer support with natural turn-taking
- Transcribe live meetings with speaker labels using Nova-3
- Analyze call center recordings for sentiment and compliance
- Generate captions for video content with low latency
- Create multilingual voice assistants with Flux conversational STT
- Automate medical transcription at scale
Models Under the Hood
as of 2026-08-14
Limitations
- Free tier limited to $200 credit; no perpetual free tier.
- Concurrency limits on lower tiers: STT up to 50 REST, 150 WSS on Pay-as-you-go, up to 225 WSS on Growth.
- Self-hosted and custom models may require Enterprise plan.
- API-first design has a learning curve for non-developers.
as of 2026-08-15
Verification history
We have re-verified Deepgram 16 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 16 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Deepgram tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Pay As You Go
$0/mo ($200 free credit)
Ideal for
Solo developer or startup exploring voice AI prototypes with a $200 free credit, needing full API access without a credit card commitment.
What this tier adds
Starting tier with a $200 free credit, all endpoints available, but limited concurrency (STT 50 REST/150 WSS, TTS 45) and no volume discounts.
Growth
$4K+/year
Ideal for
Growing applications with predictable usage that can pre-pay $4K+ annually to get up to 20% savings and higher concurrency (up to 225 WSS).
What this tier adds
Requires $4K+/year pre-payment, boosts concurrency limits (225 WSS STT, 60 TTS/Voice Agent), and offers discounted per-minute rates.
Enterprise
Contact Sales
Ideal for
Large enterprises needing custom models, self-hosted deployment, advanced security/compliance, and tailored SLAs with high volume.
What this tier adds
Adds custom model training, self-hosting, and dedicated support; pricing is contact-sales only.
Where the pricing makes sense
The company stage and team size where Deepgram's pricing actually pencils out — and where peers do it cheaper.
Deepgram's pay-as-you-go model fits developers and startups exploring voice AI, with a $200 free credit to start. For high-volume batch transcription, Nova-3 at $0.0048/min undercuts many rivals, though AssemblyAI's similar tier may edge it out on bulk pricing. The Growth tier ($4K+/yr) is best for teams already spending that much monthly—the 20% pre-paid discount is real, but the upfront commitment can sting small teams.
Setup time & first value
How long it actually takes to get something useful out of Deepgram — broken out by persona, not the marketing-page minute.
Developers: under 30 minutes to get a live transcription demo running via the Playground or SDK; a full voice agent with the Voice Agent API typically takes a day to wire up. Batch transcription: minutes to first result. Platform/enterprise: self-hosting or custom models require more time—expect days to weeks depending on infrastructure.
Switching to or from Deepgram
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From AssemblyAI: 'Deepgram supports reduced latency through the Flux models and a unified Voice Agent API, but you'll need to adapt to its pricing-per-minute model and possible concurrency caps on lower tiers.'
- →From Google Cloud STT: 'Migrating involves rewriting API calls to Deepgram's REST/WebSocket endpoints; the supported audio formats and features like Speaker Diarization are comparable.'
- ↗To AssemblyAI: 'AssemblyAI offers a similar per-minute model and add-ons; Deepgram's self-hosting and custom models may be harder to replicate.'
- ↗To Azure Speech: 'Azure provides a broader ecosystem if you're already on Microsoft, but expects you to handle turn-taking and interruption logic yourself—Deepgram's Voice Agent API does it out of the box.'
Integrations
Resources & Guides
- Documentationdeepgram.com
Welcome to Deepgram's Docs!
Deepgram developer documentation — APIs, SDKs, and tools for speech-to-text, text-to-speech, voice agents, and audio intelligence.
- Learndeepgram.com
Resources and Tools Created to Inspire
Explore resources and tools created to inspire creativity, perform deep learning, and equip your company with Voice AI best practices.
- Resourcedeepgram.com
Support
Get help with Deepgram — AI-powered support in Slack, community forums, and direct assistance
- Resourcedeepgram.com
Changelog
Helpful link from deepgram.com
- Resourcedeepgram.com
Welcome to the AI Glossary
The Deepgram AI Glossary: Your definitive resource on the world of machine learning, applied deep learning, and the field of Language AI.
- Resourcedeepgram.com
AI Minds The Podcast
Discover how the world’s top companies are building with an AI-first approach on the AI Minds podcast, brought to you by Deepgram. Listen now!
Tutorials & Learning
Tools that pair well with Deepgram
Common stack mates teams adopt alongside Deepgram, with the specific reason each pairing earns its keep.
Featured Head-to-Head Comparisons
Assemblyai vs Deepgram
If your priority is low-latency real-time STT with a unified Voice Agent API and you value TTS integration or self-hosting, Deepgram is the stronger choice. For developers who need high-accuracy batch transcription with rich speech understanding (chapters, summaries, sentiment) and an LLM gateway, AssemblyAI pulls ahead. Both are excellent, but pick based on workflow: live vs. pre-recorded.
Deepgram vs Whisper
Deepgram wins for real-time production use like voice agents and contact centers with its low-latency APIs and enterprise integrations. Whisper is ideal for budget-constrained projects needing offline multilingual transcription with zero cost. Choose based on latency needs and infrastructure support.
Alternatives to Deepgram
View allAssemblyAI
Production-grade speech-to-text and voice agent APIs for building voice AI.
ElevenLabs
ElevenLabs: AI voice platform for text-to-speech, voice cloning, dubbing, and agents
Fish Audio
Free expressive text-to-speech and voice cloning API with emotion control
Frequently Asked Questions
Best-of guides
Topics
Used Deepgram? Help shape our editorial sentiment research.


