Deepgram
Deepgram gives you speech-to-text, text-to-speech and a Voice Agent API from one vendor, priced per minute.
If you're building a real-time voice agent and don't want to glue three vendors together, Deepgram is the shortest path to a working loop. Flux STT handles turn-taking and interruptions natively, the Voice Agent API collapses STT, TTS and LLM routing into one endpoint, and published per-minute rates (streaming Nova-3 Monolingual at $0.0048/min, Flux English at $0.0065/min on Pay As You Go) make cost modelling straightforward in a category where competitors often bill in awkward increments. Against AssemblyAI and Google Cloud Speech-to-Text, the differentiator is the unified agent surface plus self-hosted deployment. It is API-first: teams wanting a hosted agent UI or a no-code builder
Verified 8d ago · liveness 95/100 · cite: rightaichoice.com/tools/deepgram
- Developers building real-time voice agents who want STT, TTS and LLM orchestration behind one endpoint
- Contact centres running live transcription, redaction and call analytics at scale
- Global products that need multiple languages handled inside a single conversation
- Enterprises with data-residency or compliance constraints that require self-hosted or custom models
- Teams that need an out-of-the-box agent UI or dashboard instead of an API
- Non-technical users who want drag-and-drop workflow builders
- Projects whose audio is so niche that off-the-shelf Nova-3 and Flux accuracy won't hold up
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Deepgram if you want a finished voice agent product with a UI and dashboards rather than an API you build on, or if your peak concurrency needs exceed the published WSS and TTS caps without moving to Enterprise.
Streaming rates shown on the pricing page are limited-time promotional prices — Nova-3 Monolingual reverts from $0.0048/min to $0.0077/min when the promotion ends, so build your model on the regular price.
Pay As You Go suits solo developers and seed-stage startups — no minimums, no expiration, $200 credit to start. Growth at $4K+/year in pre-paid credits (saving up to 20%) fits Series A to mid-market teams with predictable monthly audio volume. Enterprise is custom-priced for large volumes, data residency and self-hosting. Against AssemblyAI and Google Cloud Speech-to-Text, Deepgram competes on published minute-level rates and concurrency rather than headline discounts.
In short
Deepgram — Deepgram gives you speech-to-text, text-to-speech and a Voice Agent API from one vendor, priced per minute. Best for Developers building real-time voice agents who want STT, TTS and LLM orchestration behind one endpoint, Contact centres running live transcription, redaction and call analytics at scale, Global products that need multiple languages handled inside a single conversation. Plans from $4/yr.
What's new in Deepgram
Checked 8 days agoAcross the latest 5 updates: 2 feature updates, 1 launch and 2 changelog entries.
Toggle Numerals Mid-Stream on Flux STT
Flux STT can now switch to digits for a PIN, phone number or order number and back to words without reconnecting, by sending numerals as a boolean in a Configure message. This replaces the connection-time-only behaviour from the July 17 entry.
Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese
Improved Nova-3 monolingual models released for nine existing languages, enhancing transcription quality for both batch and streaming workloads.
Correction: Browser Agent SDK has no client-side VAD
Deepgram corrected documentation that described the Browser Agent SDK as shipping client-side Silero voice activity detection. The vad option, speechThreshold, silenceThreshold, VAD events and peer dependencies never existed in any released package.
Correction: token minting uses ttl_seconds
A server-side token-minting example on the Browser Agent SDK overview posted ttl instead of ttl_seconds. The API accepted the unrecognised field and silently issued a 30-second default token, so requests built from the old example expired early.
Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)
nova-3-pharma is a new Nova-3 model purpose-built for pharmaceutical vocabulary with a focus on accurate drug-name recognition, designed for pharmacy and healthcare voice-agent workflows. Available in English for both batch and streaming.
What people actually say about Deepgram — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
39 mentions across 4 sources (Hacker News, Product Hunt, Stack Overflow, Lemmy) · researched Aug 18, 2026.
Average across the 4 sources that answered — each source counts once, not each post.
- +Low latency for real-time voice agents (community mentions).
- +Unified Voice Agent API simplifies STT+TTS+LLM integration.
- +High accuracy with Nova-3 models, especially multilingual.
- +Flexible deployment: cloud or self-hosted.
- +Generous free tier to test and prototype (community acknowledges).
- −Self-hosting setup can be complex and requires resources.
- −Free tier limits may surprise high-volume users.
- −Cloud dependency undermines 'local-first' claims.
- −Documentation could be clearer for beginners (async examples).
- −Advanced features like diarization may be pricey.
- • Add-ons like speaker diarization may incur extra per-minute charges
- • Self-hosting requires significant infrastructure costs
- • Voice Agent API might charge per call hour, which can add up
Viability Score
How well maintained and how widely used is Deepgram? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Real-time streaming speech-to-text with Flux and Nova-3 models
- Flux STT conversational transcription with built-in turn detection and interruption handling
- Flux Multilingual recognises multiple languages inside a single conversation
- Batch pre-recorded transcription for archived audio
- Text-to-speech with Flux TTS, Aura-2 and Aura-1 voices
- Unified Voice Agent API combining STT, LLM orchestration and TTS in one call
- Runtime turn-taking control: override, suppress or blend end-of-turn detection
- Nova-3 Monolingual and Nova-3 Multilingual with automatic language detection across 45+ languages
- Nova-3 Pharma model for pharmaceutical and drug-name recognition (English)
- Speaker Diarization for multi-speaker detection
- Audio Intelligence API for emotion and sentiment analysis
- Redaction of PII such as social security numbers, credit cards and phone numbers
- Keyterm Prompting to boost accuracy on domain jargon, product names and acronyms
- Entity Detection to extract structured data, plus Smart Formatting and Numerals
- Custom speech-to-text model training on proprietary datasets
About Deepgram
Deepgram is a voice AI platform built for developers, platforms and enterprises that need speech-to-text, text-to-speech and conversational agents without stitching three vendors together. The centrepiece is a single Voice Agent API: speech-to-text, LLM orchestration and text-to-speech sit behind one call instead of three integrations, which cuts latency and the integration work that usually eats a sprint. The current generation is Flux. Flux STT is conversation-aware transcription for real-time agents with built-in turn detection and interruption handling rather than bolted-on add-ons, and Flux Multilingual handles several languages inside one conversation. Flux TTS, announced in September 2026, reads the whole conversation rather than the current line and holds tone and context across turns. Nova-3 remains the workhorse transcription family, with monolingual and multilingual variants across 45+ languages, plus a September 2026 Nova-3 Pharma model purpose-built for drug-name recognition in pharmacy and healthcare workflows. Text-to-speech also runs on Aura-2 and Aura-1. Beyond raw transcription, Audio Intelligence adds emotion and sentiment analysis, and add-ons like Redaction, Keyterm Prompting, Entity Detection, Speaker Diarization, Smart Formatting and Numerals turn a transcript into something downstream systems can act on. Custom models can be trained on proprietary datasets for edge cases. Deployment is cloud or self-hosted, with WebSocket and REST APIs plus Python, JavaScript, Go, .NET and Java SDKs, and a self-hosted release train on quay.io container images. Pricing is usage-based: $200 in free credit, then published per-minute rates such as $0.0048/min streaming Nova-3 Monolingual and $0.0065/min streaming Flux English on Pay As You Go. AssemblyAI and Google Cloud Speech-to-Text are the usual comparisons; Deepgram's pitch is low latency, one API surface and minute-level published rates.
Behind the Verdict
Deepgram's strength is that it stopped pretending transcription is the whole product. The Voice Agent API is the real pitch: one call carries user audio through STT, LLM orchestration and TTS, with business logic and external systems hooked in at the same layer. That removes the two integrations you'd otherwise own and maintain, and it removes the latency those hops add. Flux is the current generation and the most interesting part of the stack. Flux STT is built for conversational audio with turn detection and natural interruption handling as first-class behaviour rather than a threshold you tune yourself, and Flux Multilingual recognises languages within a single conversation — genuinely useful for support lines that mix languages mid-call. Flux TTS, launched September 2026, reads the whole conversation rather than the current sentence, so tone and expression carry across turns instead of resetting each line. Nova-3 is the volume workhorse: monolingual and multilingual variants over 45+ languages, with automatic language detection, Smart Formatting, Speaker Diarization and Keyterm Prompting. Deepgram keeps shipping language improvements — September 2026 releases improved Nova-3 monolingual models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu and Vietnamese on one date, and Flemish, German (Switzerland), Lithuanian and Portuguese on another. Nova-3 Pharma, also September 2026, targets drug-name accuracy for pharmacy and healthcare voice agents. Audio Intelligence layers emotion and sentiment on top, and Redaction, Entity Detection and Numerals make transcripts machine-usable. Operationally, Deepgram is more transparent than most of this category. Rates are published per minute with the promotional and regular price shown side by side, concurrency limits are stated per plan per API, and self-hosted deployments ship on a dated container release train with an explicit minimum NVIDIA driver version. SDK coverage spans Python, JavaScript, Go, .NET and Java, plus React bindings and a Browser Agent SDK. The honest friction is in the API-first design and the concurrency model. Pay As You Go caps STT at up to 50 REST and 150 WSS connections, TTS and Voice Agent at 45, and Audio Intelligence at 10; Growth lifts WSS to 225 and TTS/Voice Agent to 60 but still leaves REST at 50. Deepgram Whisper Cloud is capped at 5 concurrent connections on both plans. Flux TTS promotional access is 45 concurrent streaming connections in the US and only 5 in the EU and AU. If you're sizing a contact centre, those ceilings matter more than the per-minute rate. The other caveat is documentation churn: Deepgram published two corrections in September 2026, one retracting client-side VAD claims for the Browser Agent SDK and one fixing a token-minting example that used ttl instead of ttl_seconds and silently issued 30-second tokens. Worth reading the changelog before you copy an example. Where it fits: product teams embedding voice into an
Researching Deepgram? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Deepgram actually fits — and what changes day-one when you adopt it.
You're wiring a support agent into an existing app and don't want three vendor SDKs. You point the Voice Agent API at your LLM, set turn-taking to blend with Flux STT's end-of-turn detection, and pick a Flux TTS voice. Business logic and CRM calls hook in at the orchestration layer.
Outcome: A working real-time loop with interruptions and barge-in handled by the platform instead of your state machine, priced at $0.0065/min streaming Flux English on Pay As You Go.
You need live transcription of inbound calls with PII stripped before anything hits storage. You stream through Nova-3 Monolingual, turn on Redaction for SSNs and card numbers, and run Audio Intelligence over the recorded audio for sentiment.
Outcome: Redacted transcripts plus sentiment scores, budgeted at $0.0048/min transcription plus $0.0020/min Redaction plus Audio Intelligence on Pay As You Go — with the 10-connection Audio Intelligence REST cap as the thing to size against.
Your data can't leave the region. You deploy the self-hosted API and Engine images from quay.io, verify the NVIDIA driver meets the >=580 open-kernel requirement, and use the India endpoint (api.in.deepgram.com) for the traffic that can stay in-country.
Outcome: Deepgram inference inside your own infrastructure on a dated release train, with custom models trained on your proprietary audio where off-the-shelf accuracy isn't enough.
Use Cases
- Build real-time voice agents for customer support with natural turn-taking and interruption handling
- Transcribe live meetings with speaker labels using Nova-3 and Speaker Diarization
- Analyse call centre recordings for sentiment and compliance with Audio Intelligence
- Generate low-latency captions for video content
- Create multilingual voice assistants that switch languages mid-conversation with Flux Multilingual
- Run pharmacy or healthcare voice agents with Nova-3 Pharma's drug-name accuracy
- Self-host STT and TTS inside your own infrastructure for data-residency requirements
Models Under the Hood
as of 2026-10-09
Limitations
- Developer/API-first platform: STT concurrency is capped at up to 50 for the REST API on Pay As You Go and Growth, and up to 150 WSS on Pay As You Go rising to 225 on Growth.
- Deepgram Whisper Cloud is limited to 5 concurrent connections on both plans.
- Text-to-speech and Voice Agent are capped at 45 concurrent REST/WSS connections on Pay As You Go and 60 on Growth; Audio Intelligence is 10 on both.
- Flux TTS promotional access is capped at 45 concurrent streaming connections in the US and 5 in the EU and AU, and the free option is a $200 credit rather than a perpetual free tier (no credit card required to start).
- Documentation has needed correction: a September 2026 changelog entry retracted client-side Silero VAD claims for the Browser Agent SDK, and another fixed a token-minting example that used ttl instead of ttl_seconds and silently issued 30-second tokens.
as of 2026-09-30
Verification history
We have re-verified Deepgram 19 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 19 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Annual billing (Growth) runs $4/yr — about $2,396 less than paying monthly for a year.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Deepgram tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Pay As You Go
$200 free credit, then pay-as-you-go
Ideal for
Solo developers and seed-stage teams prototyping voice features without a contract — $200 credit to start, no minimums and no expiration.
What this tier adds
Starting tier: pay per minute with no commitment, concurrency up to 50 REST / 150 WSS for STT, 45 for TTS and Voice Agent, 10 for Audio Intelligence.
Growth
$4K+/year (pre-paid credits, save up to 20%)
Ideal for
Series A to mid-market teams with predictable monthly audio volume who can commit $4K+/year in pre-paid credits.
What this tier adds
Adds up to 20% savings on pre-paid credits, higher WSS concurrency (225 STT, 60 TTS/Voice Agent) and discounted Redaction and Keyterm Prompting — REST concurrency and Audio Intelligence stay at 50 and 10.
Enterprise
Custom
Ideal for
Large-volume buyers with data-residency, self-hosting or custom-model requirements, or partners embedding voice into their own platform.
What this tier adds
Adds custom speech-to-text models trained on proprietary datasets, self-hosted deployment, partner embedding programs and dedicated support.
Where the pricing makes sense
The company stage and team size where Deepgram's pricing actually pencils out — and where peers do it cheaper.
Pay As You Go suits solo developers and seed-stage startups — no minimums, no expiration, $200 credit to start. Growth at $4K+/year in pre-paid credits (saving up to 20%) fits Series A to mid-market teams with predictable monthly audio volume. Enterprise is custom-priced for large volumes, data residency and self-hosting. Against AssemblyAI and Google Cloud Speech-to-Text, Deepgram competes on published minute-level rates and concurrency rather than headline discounts.
Setup time & first value
How long it actually takes to get something useful out of Deepgram — broken out by persona, not the marketing-page minute.
Grab a free API key and you can make a first streaming request in well under an hour — SDKs exist for Python, JavaScript, Go, .NET and Java, and there's a Playground for testing models without writing code. A full Voice Agent loop wired to your own LLM and business logic is a day or two. Self-hosted deployment is the long pole: pulling the API, Engine and license-proxy containers and meeting the
Switching to or from Deepgram
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From AssemblyAI: swap the streaming endpoint for Deepgram's WebSocket API, map your model choice to Nova-3 Monolingual or Flux English, and re-add features you used there as Deepgram add-ons such as Redaction and
- →From Google Cloud Speech-to-Text: replace the gRPC client with a Deepgram SDK, move recognition config to model and language parameters, and re-implement diarization and formatting via Deepgram's built-in options.
- →From OpenAI Whisper self-hosted: move batch workloads to Deepgram's pre-recorded Nova-3 endpoint, or keep Whisper Large through Deepgram's pre-recorded Whisper Large model at $0.0048/min.
- →From a hand-rolled STT+LLM+TTS pipeline: collapse the three hops into the Voice Agent API and set turn-taking at runtime instead of tuning silence thresholds yourself.
- ↗To AssemblyAI: map Nova-3 to Universal-2 style models and re-implement Redaction, Keyterm Prompting and Audio Intelligence as equivalent features there.
- ↗To Google Cloud Speech-to-Text: move streaming to the gRPC streaming API and rebuild diarization and custom vocabulary through Google's recognition config.
- ↗To a managed agent platform: replace the Voice Agent API with a hosted agent product if you decide you want a UI and dashboards rather than an API to build on.
- ↗To self-hosted open models: swap Deepgram cloud for a local Whisper or equivalent deployment, and move turn detection and interruption handling back into your own code.
Integrations
Resources & Guides
- Documentationdeepgram.com
Welcome to Deepgram's Docs!
Deepgram developer documentation — APIs, SDKs, and tools for speech-to-text, text-to-speech, voice agents, and audio intelligence.
- Learndeepgram.com
Resources and Tools Created to Inspire
Explore resources and tools created to inspire creativity, perform deep learning, and equip your company with Voice AI best practices.
- Resourcedeepgram.com
Support
Get help with Deepgram — AI-powered support in Slack, community forums, and direct assistance
- Resourcedeepgram.com
Changelog
Helpful link from deepgram.com
- Resourcedeepgram.com
Welcome to the AI Glossary
The Deepgram AI Glossary: Your definitive resource on the world of machine learning, applied deep learning, and the field of Language AI.
- Resourcedeepgram.com
AI Minds The Podcast
Discover how the world’s top companies are building with an AI-first approach on the AI Minds podcast, brought to you by Deepgram. Listen now!
Tutorials & Learning
YouTube returned 6 videos for “Deepgram”, and we withheld 6: 6 could not be judged, because “Deepgram” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Deepgram.
Official links
Tools that pair well with Deepgram
Common stack mates teams adopt alongside Deepgram, with the specific reason each pairing earns its keep.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
ElevenLabs
ElevenLabs turns text into ultra-realistic AI voice, music, dubbing and conversational voice agents from one credit pool.
AssemblyAI
Voice AI infrastructure for developers: speech-to-text, speech understanding, guardrails, and LLM routing on one API key.
Featured Head-to-Head Comparisons
Alternatives to Deepgram
View allFish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
ElevenLabs
ElevenLabs turns text into ultra-realistic AI voice, music, dubbing and conversational voice agents from one credit pool.
AssemblyAI
Voice AI infrastructure for developers: speech-to-text, speech understanding, guardrails, and LLM routing on one API key.
Frequently Asked Questions
Best-of guides
Topics
Used Deepgram? Help shape our editorial sentiment research.