Speechmatics
Multilingual real-time speech-to-text API with sub-second latency and enterprise-grade security.
Speechmatics is a strong pick for enterprises needing compliant, high-accuracy multilingual ASR with low latency. The Melia model, on-device integrations (Adobe Premiere, Stenograph), and credit-based pricing are standout features, though smaller teams may find the contact-sales enterprise tier limiting. For teams that need sub-second multilingual transcription with strong security, it's a solid choice compared to generic cloud STT providers.
Verified 5d ago · liveness 82/100 · cite: rightaichoice.com/tools/speechmatics
- Developers building voice agents or real-time transcription apps needing multilingual support and sub-second latency
- Healthcare organizations requiring HIPAA-compliant, accurate medical transcription with specialized vocabulary
- Media and broadcast teams needing live captioning for events, sports, or news with high accuracy
- Contact centers seeking real-time call analytics and agent assist with enterprise compliance
- Non-technical users wanting a no-code transcription solution with drag-and-drop interface
- Hobbyists or very small projects without budget for Pro or Enterprise tiers (free credit is limited)
- Use cases requiring out-of-the-box integrations with niche CRMs or legacy systems
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Speechmatics if you're a non-technical user wanting a no-code drag-and-drop transcription tool, or if you need multilingual text-to-speech (currently English-only).
Going over 500 hours per month on the Pro tier doesn't automatically apply the 20% discount until you contact sales; you must negotiate for it.
Speechmatics' credit-based pricing scales well for enterprises with volume discounts (20% over 500 hr/month) and custom enterprise plans, but the free tier ($100 credit) is limited. Compared to AWS Transcribe (pay-as-you-go) and Google STT, Speechmatics may cost more per hour for batch (Standard $0.24/hr vs AWS ~$0.024/min), but it offers better multilingual accuracy and security. Best for mid-to-large teams with budget for Pro or Enterprise; small projects may find free credit too limiting.
In short
Speechmatics — Multilingual real-time speech-to-text API with sub-second latency and enterprise-grade security. Best for Developers building voice agents or real-time transcription apps needing multilingual support and sub-second latency, Healthcare organizations requiring HIPAA-compliant, accurate medical transcription with specialized vocabulary, Media and broadcast teams needing live captioning for events, sports, or news with high accuracy. Free to start; paid plans from $0.129/mo.
What's new in Speechmatics
Checked 5 days agoAcross the latest 5 updates: 3 feature updates, 1 launch and 1 pricing change.
Speechmatics on Zapier: No-Code Speech-to-Text Automation
Speechmatics integrates with Zapier, enabling no-code speech-to-text automation workflows for non-technical users.
Speechmatics launches Medical Model for real-time clinical transcription
New medical-specific speech-to-text model optimized for real-time clinical transcription, cutting errors on key terms by up to 50%.
Speaker Focus: Fixing Voice AI for the real world
Introduces Speaker Focus feature to improve voice AI accuracy in real-world, multi-speaker scenarios.
Stenograph and Speechmatics Announce Industry-First On-Device Integration for CATalyst VP
Speechmatics partners with Stenograph for on-device integration into CATalyst VP, enabling offline transcription for legal professionals.
A Simpler Way to Pay: Speechmatics Is Moving to Credits
Speechmatics shifts from per-minute pricing to a credit-based system, simplifying usage costs.
What people actually say about Speechmatics — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
26 mentions across 2 sources (Hacker News, YouTube) · researched Aug 19, 2026.
- +Excellent multilingual accuracy with support for 56+ languages and code-switching
- +Reliable speaker diarization, ideal for multi-speaker conversations
- +Military-grade security: ISO 27001, HIPAA, SOC 2 Type II compliant
- +Zero data logging by default ensures privacy for regulated industries
- +Sub-second latency for real-time streaming via WebSocket and REST API
- −Priced higher than competitors like Voxtral and Deepgram, at $0.004/min
- −No free tier substantial enough for large-scale testing
- −Some users report terrible customer support experience, though not universally
- −On-device deployment requires integration with specific partners, not standalone
- −Steep learning curve for non-developers due to heavy API focus
- • Enterprise features like custom vocabulary and on-prem deployment likely require higher-tier plans
- • Volume discounts require contacting sales, no transparent pricing
Viability Score
How well maintained and how widely used is Speechmatics? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: August 2026
How we score →Key Features
- Real-time speech-to-text with sub-second latency
- Batch transcription with microbatching
- 55+ languages and dialects
- Melia multilingual model with code-switching
- Speaker diarization and speaker focus
- Custom vocabulary and custom language models
- Medical Model for clinical transcription
- Low-latency text-to-speech (English, more coming)
- Translation, summaries, chapters, sentiment, topics bolt-ons
- On-device deployment (Adobe Premiere, Stenograph CATalyst VP)
- On-premise, private cloud, container, virtual appliance deployment
- Zero data logging by default
- ISO 27001, HIPAA, SOC 2 Type II compliance
- Real-time streaming via WebSocket and REST API
- Voice agent API with speaker-aware STT and TTS
About Speechmatics
Speechmatics is a speech recognition platform offering real-time and batch speech-to-text, text-to-speech, and voice agent APIs for developers and enterprises. With 55+ languages and dialects covering over 4 billion people, it delivers STT in less than a second, making it a strong choice for live captioning, voice agents, contact center analytics, medical transcription, and legal reporting. Its core differentiator is the Melia multilingual model, now in production preview for batch transcription, which supports code-switching across 55+ languages—so speakers can switch mid-sentence and still get accurate transcripts. Speechmatics includes a Medical Model for real-time clinical transcription, cutting errors on key terms by up to 50% and detecting health signals from short voice clips. For legal professionals, on-device deployment through Stenograph's CATalyst VP brings industry-first on-prem speech recognition to court reporting. Recent updates include Zapier integration for no-code automation and a new Speaker Focus feature for improved multi-speaker accuracy. Developers can take advantage of real-time streaming via WebSocket, REST, and native integrations like LiveKit, plus features such as speaker diarization, custom vocabulary, language identification, and precise timestamps. Security is built-in: the platform is ISO 27001, HIPAA, and SOC 2 Type II compliant, with zero data logging by default. Flexible deployment options—cloud, on-premise, private cloud, container, virtual appliance, and on-device—give privacy-critical industries control over their data. Pricing is credit-based, starting with $100 free credit, and scales with volume discounts (20% off over 500 hours per month) and enterprise custom pricing. If you need high-accuracy, multilingual ASR with enterprise-grade security and low latency, Speechmatics competes directly with AWS Transcribe and Google STT, but stands apart with its focus on real-world accents, code-switching, and compliant deployment flexibility. For teams building voice AI at scale, Speechmatics is a dependable API-first choice.
Behind the Verdict
Speechmatics stands out in the crowded speech-to-text market by focusing on three things: accuracy across real-world accents and code-switching, enterprise-grade compliance and security, and flexible deployment options. The Melia model is a differentiator—it handles 55+ languages and lets speakers switch mid-sentence without losing accuracy. The recent launch of the Medical Model (2026-07-29) makes it especially attractive for healthcare applications, where accuracy on specialized terminology is critical. The Stenograph partnership (2026-07-22) offers true on-device capability for legal professionals, which is a niche but compelling use case. However, the pricing model, while simplified with credits, can still be complex for small teams. The free tier gives you $100 in credit, which is generous but limited; you'll need to move to Pro or Enterprise for substantial usage. The Pro tier starts at $0.129/hr, which is competitive but not the cheapest. If you're a hobbyist or a very small project, you might find the free credit insufficient, and the jump to paid tiers might be more than you need. For enterprises, Speechmatics is a solid choice. The security certifications (ISO 27001, HIPAA, SOC 2 Type II) and zero data logging are reassuring. The ability to deploy on-premises or on-device is a major plus for privacy-critical industries. The 20% volume discount over 500 hours/month helps at scale. In terms of limitations, Text-to-Speech is currently English-only, which is a gap if you need multilingual TTS. There is no no-code drag-and-drop interface; you'll need some technical know-how to integrate the API. Compared to alternatives like AWS Transcribe or Google STT, Speechmatics offers better multilingual support and code-switching, plus more flexible deployment. However, those cloud giants may be cheaper for simple use cases. Speechmatics is best for teams that need high accuracy on real-world speech and prioritize security and compliance.
Researching Speechmatics? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Speechmatics actually fits — and what changes day-one when you adopt it.
Using the Voice Agent API with LiveKit integration
Outcome: Deploy a multi-turn voice agent in 55+ languages with speaker-aware STT and low-latency TTS, achieving sub-second response times.
Implementing ambient medical scribes with the Medical Model
Outcome: Reduce documentation time by up to 50% with real-time clinical transcription, improving accuracy on key medical terms and reducing clinician burnout.
Analyzing call recordings for insights
Outcome: Automatically transcribe calls in real-time, with sentiment and topic detection, to improve agent performance and customer satisfaction.
Use Cases
- Transcribe live sports events with real-time captions at scale.
- Reduce documentation time in hospitals with ambient medical scribes using Medical Model.
- Empower voice agents to handle multi-turn, multi-speaker conversations in 55+ languages.
- Analyze call center recordings to extract insights and improve agent performance.
- Caption courtroom proceedings with high accuracy across diverse accents.
- Integrate on-device speech recognition into video editing software for offline captioning.
- Set up no-code speech-to-text automation via Zapier integrations.
- Build meeting platforms with automated note-taking in multiple languages.
Models Under the Hood
as of 2026-08-31
Limitations
- Speech-to-Text supports 55+ languages and dialects, while Text-to-Speech is currently English-only with more languages coming soon.
- The free tier includes $100 in credit and 2 concurrent real-time sessions, while the Pro tier offers 50 concurrent real-time sessions and 10 file jobs per second.
- Enterprise offers unlimited scale with no rate limits.
- Pricing has shifted to a credit-based system, with Pro pricing starting from $0.129/hr with a 20% discount available.
as of 2026-08-28
Verification history
We have re-verified Speechmatics 16 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 16 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Speechmatics tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0
Ideal for
Developers and early-stage projects exploring speech-to-text capabilities; includes $100 credit and 2 concurrent real-time sessions for prototyping.
What this tier adds
No credit card required and $100 free credit to get started; enough for initial testing and small demos.
Pro
from $0.129/hr
Ideal for
Growing production use with higher concurrency needs (50 concurrent real-time sessions) and 20% discount over 500 hours/month.
What this tier adds
Adds a card for billing, raises concurrency to 50, increases batch jobs to 10/sec, and includes online email support.
Enterprise
Custom
Ideal for
Large enterprises needing unlimited scale, custom models, on-premise or private cloud deployment, and dedicated support.
What this tier adds
Contact sales for volume discounts, no rate limits, and access to custom language and voice development plus prioritized service.
Where the pricing makes sense
The company stage and team size where Speechmatics's pricing actually pencils out — and where peers do it cheaper.
Speechmatics' credit-based pricing scales well for enterprises with volume discounts (20% over 500 hr/month) and custom enterprise plans, but the free tier ($100 credit) is limited. Compared to AWS Transcribe (pay-as-you-go) and Google STT, Speechmatics may cost more per hour for batch (Standard $0.24/hr vs AWS ~$0.024/min), but it offers better multilingual accuracy and security. Best for mid-to-large teams with budget for Pro or Enterprise; small projects may find free credit too limiting.
Setup time & first value
How long it actually takes to get something useful out of Speechmatics — broken out by persona, not the marketing-page minute.
For developers, you can get started with the free tier and API access in under 5 minutes, with sample code and docs. Adding integrations like LiveKit takes about 30 minutes. For enterprise deployments, allow 2-4 weeks for custom models and on-premise setup, depending on complexity.
Switching to or from Speechmatics
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From AWS Transcribe: Re-architect your API calls to Speechmatics endpoints, but the transition is straightforward due to similar REST/WebSocket patterns.
- →From Google STT: Adapt your code to use Speechmatics' API; focus on multilingual features and custom vocabularies.
- ↗To AWS Transcribe: Switch API calls to AWS' Transcribe service; compare pricing and features to ensure fit.
- ↗To Google STT: Use Google's speech-to-text API, but note that multilingual code-switching may be less robust.
Integrations
Resources & Guides
- Resourcespeechmatics.com
Blog & Latest Speech Recognition News
Stay up-to-date with the very latest product news, company news and artificial intelligence (AI) research. Bookmark our technology blog today!
- Resourcespeechmatics.com
AI Speech Technology | Speech APIs powering Voice AI
Speechmatics offer the most accurate AI speech technology for enterprise - with AI transcription, real-time translation and text-to-speech components. Try our Speech API today!
- Resourcespeechmatics.com
About Us
Speechmatics: Leading provider of automatic speech recognition technology. Unlock audio and video value with innovative ASR solutions. Transform spoken language into written text and revolutionize speech processing.
- Resourcegithub.com
Speechmatics
Speechmatics has 52 repositories available. Follow their code on GitHub.
Tutorials & Learning
Official links
Tools that pair well with Speechmatics
Common stack mates teams adopt alongside Speechmatics, with the specific reason each pairing earns its keep.
Alternatives to Speechmatics
View allFrequently Asked Questions
Best-of guides
Used Speechmatics? Help shape our editorial sentiment research.


