Typecast AI
Expressive text-to-speech API with Smart Emotion, 600+ voices in 35+ languages, and 170-210ms streaming for real-time agents.
If your agent needs to sound like it means what it says, Typecast's Smart Emotion plus 170-210ms streaming is the combination most competitors make you assemble by hand, and the September 2026 remove_silence_ms parameter shows the team is still shipping developer-facing polish rather than voice-count marketing. Against OpenAI's or Azure's TTS you get more explicit emotion control and a lower published unit cost; against ElevenLabs you give up some brand recognition and the on-prem option, and you cannot buy a plan that follows you across iOS, Android and web. Evaluation is genuinely free, but commercial rights only start at $5/mo billed annually ($54/yr) on Basic, and long-form audio still
Verified 3d ago · liveness 79/100 · cite: rightaichoice.com/tools/typecast-ai
- Developers building voice agents that need sub-200ms streaming TTS
- Content teams automating YouTube, TikTok and ad voiceovers without a studio
- Audiobook and podcast producers who need expressive narration with emotion control
- Product teams adding multilingual narration to an existing app
- Teams that require on-premise or air-gapped deployment
- Long-form batch rendering where per-character credit cost dominates the budget
- Projects needing unlimited commercial output on a zero budget
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Typecast if you need on-premise or air-gapped deployment, or if you are optimizing purely on per-character cost at very large batch volume.
Credits are only consumed when you download audio — generation and playback are free — so budgeting by trial usage will understate your real bill.
For API work, Typecast's Lite tier at $15/mo (200k credits, 5 concurrent requests, 50 cloning slots) undercuts ElevenLabs' comparable entry tiers and sits well below Azure or Google enterprise TTS once you factor in the emotion layer you would otherwise build. The web studio runs $5/mo Basic (billed yearly at $54) up to $69/mo Business ($744/yr) for teams. Anyone past a few million characters a month should be pricing Enterprise rather than stacking Plus overage.
In short
Typecast AI — Expressive text-to-speech API with Smart Emotion, 600+ voices in 35+ languages, and 170-210ms streaming for real-time agents. Best for Developers building voice agents that need sub-200ms streaming TTS, Content teams automating YouTube, TikTok and ad voiceovers without a studio, Audiobook and podcast producers who need expressive narration with emotion control. Free to start; paid plans from $5/mo.
What's new in Typecast AI
Checked 3 days agoAcross the latest 5 updates: 3 feature updates, 1 launch and 1 changelog entry.
Typecast adds silence control to official SDKs and integrations
Official SDKs, Cast CLI, n8n, Zapier, Pipecat and the hosted Typecast API MCP now support remove_silence_ms, letting developers cap how much silence is retained in generated speech.
Typecast API adds remove_silence_ms TTS parameter
A new remove_silence_ms parameter (0-1000ms) shortens detected silence across TTS, streaming TTS, timestamped TTS and Compose. Explicit pause segments are preserved and timestamps realign to the trimmed audio.
Typecast launches Professional Voice Cloning API and V3 voices
An async POST /v1/custom-voices/professional-clone endpoint accepts a WAV or MP3 sample plus language and returns a custom voice with completed/failed status. Ships with updated V3 voice APIs, SDK packages and Cast CLI v1.0.9.
Voice recommendations and Compose TTS endpoints released
GET /v1/voices/recommendations matches voices to a natural-language description, and POST /v1/text-to-speech/compose synthesizes multiple segments in one request with per-segment voice and speech settings.
Instant Voice Cloning endpoint released; old clone endpoint deprecated
POST /v1/voices/clone creates a voice from a WAV or MP3 sample. The original endpoint is deprecated in favor of POST /v1/custom-voices/instant-clone for new integrations.
What people actually say about Typecast AI — is it worth it?
We scanned public community sources for Typecast AI on Sep 24, 2026 and could not establish that the discussion we found is about this tool rather than something else sharing its name. Only 1 of the posts we fetched could be positively tied to Typecast AI. Rather than publish a sentiment score built on the wrong subject, we publish nothing here and re-run the scan.
Viability Score
How well maintained and how widely used is Typecast AI? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Text-to-speech API (POST /v1/text-to-speech) returning WAV or MP3
- Streaming TTS endpoint with 170-210ms time-to-first-byte for real-time agents
- Timestamp TTS with word- and character-level alignment for subtitles and karaoke
- Compose TTS endpoint synthesizing multiple segments in one request with per-segment settings
- Smart Emotion applies tone automatically from surrounding context (previous_text / next_text)
- 7 manual emotion presets plus custom emotion, intonation, pitch and speed controls
- remove_silence_ms parameter (0-1000ms) to trim detected silence without cutting explicit pauses
- Instant voice cloning from a short WAV or MP3 sample
- Professional Voice Cloning API (async) for high-accuracy intonation and tone matching
- Voice recommendations endpoint matching voices to a natural-language description
- 600+ voices across ages, tones and personalities, 35+ languages with native-level naturalness in 6
- Official SDKs for Python, JavaScript/TypeScript, Go, Rust, C#, Java, Kotlin, C, Swift, Zig, PHP, Dart and Ruby
- Cast CLI (v1.0.9) for command-line generation and agent workflows
- No-code integrations via Zapier, Make, n8n and Google Sheets
- Agent integrations including MCP server, OpenClaw, Claude Skills and Pipecat
About Typecast AI
Typecast is a cloud text-to-speech API from Neosapience, aimed at developers building conversational agents, content teams automating voiceovers, and product teams adding narration to apps. You send text to a REST endpoint and get back WAV or MP3 audio, with official SDKs covering Python, JavaScript, Go, Rust, C#, Java, Kotlin, C, Swift, Zig, PHP, Dart and Ruby, a Cast CLI (v1.0.9 as of the August 2026 release), and no-code hooks into Zapier, Make, n8n and Google Sheets if you would rather not write glue code. The pitch rests on expressiveness rather than raw voice count, though the count is real: 600+ voices spanning 35+ languages, with native-level naturalness claimed in 6, English included. Smart Emotion reads surrounding context — you can pass previous_text and next_text — and applies tone automatically; 7 manual emotion presets plus custom emotion, intonation, pitch and speed controls let you override when you want a specific read. Voice cloning comes in two flavors: Instant Cloning from a short sample, and Professional Cloning (launched August 2026 as an async API endpoint) that chases the original speaker's intonation and tone. The underlying model is Typecast's own SSFM line, with ssfm-v30 adding context-based emotion control and ssfm-v21 still available. For real-time work there is a Streaming TTS endpoint reporting 170-210ms time-to-first-byte depending on model, which is the number that matters if you are wiring voice into a live agent loop rather than batch-rendering files. A Timestamp endpoint returns word- and character-level alignments for subtitles, karaoke and lip-sync, and the July 2026 'Compose' endpoint synthesizes multiple segments in one request with per-segment voice and speech settings. The September 2026 release added a remove_silence_ms parameter (0-1000ms) across TTS, streaming TTS, timestamped TTS and Compose, now propagated to the official SDKs, Cast CLI, n8n, Zapier, Pipecat and the hosted Typecast API MCP server. Positioning is straightforward: cheaper and more curated than a global TTS API, more turnkey than assembling XTTS v2 yourself on your own GPUs. If you need on-prem or air-gapped deployment, or you are optimizing purely on per-character cost at very large volume, look elsewhere first.
Behind the Verdict
Typecast's real differentiation is not the 600-voice library — it is that emotion is a first-class API parameter rather than something you bolt on. Smart Emotion inspects the surrounding text (you literally pass previous_text and next_text) and picks the tone, and when you need a specific read you drop to one of 7 manual presets or set custom emotion, intonation, pitch and speed. For conversational agents this matters more than voice variety, because a chatbot that reads every line flatly sounds broken no matter how good the timbre is. The developer surface is unusually broad for a company this size: REST endpoints for full TTS, streaming TTS, timestamped TTS, Compose (multi-segment in one request), instant cloning, professional cloning, and voice recommendations from a natural-language description (GET /v1/voices/recommendations, July 2026). Official SDKs span Python, JS/TS, Go, Rust, C#, Java, Kotlin, C, Swift, Zig, PHP, Dart and Ruby — a list most TTS vendors do not come close to — plus a Cast CLI and a docs site that ships an llms.txt so coding agents can read it directly. MCP, Claude Skills, OpenClaw and Pipecat integrations put it in the agent stack rather than beside it. Where it gets thinner: the model is ssfm-v30 (or ssfm-v21), and Typecast does not publish concurrency-under-load latency figures, only the 170-210ms TTFB range. There is no on-prem or air-gapped deployment. The credit model is character-based, so a 10-hour audiobook at 1 credit per character is a real budget line, and Enterprise is custom-quoted with no published number. Longer-term it is a Neosapience product, so you are betting on a specialist rather than a hyperscaler whose TTS ships as one line item among hundreds. Where it fits: voice agents, AI tutors, IVR/AICC, subtitled video pipelines, and any team that wants expressive narration without hiring a studio or babysitting a GPU. Where it does not: teams with a strict data-residency requirement, anyone whose only metric is dollars per million characters at massive scale, and buyers who expect one subscription to cover web plus both mobile apps.
Researching Typecast AI? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Typecast AI actually fits — and what changes day-one when you adopt it.
You point the Streaming TTS endpoint at your LLM's output stream, set emotion_type to smart, and pass previous_text so the tone carries across turns.
Outcome: Your agent starts speaking within 170-210ms of the first chunk and sounds like it means what it says, without a separate emotion-classification step.
You trigger Typecast from Zapier or Make with a script row from Google Sheets, then pull the timestamped TTS output into your captioning step.
Outcome: Voiceovers and word-level caption timings arrive in one pipeline, and remove_silence_ms trims the dead air without cutting the deliberate pauses.
You clone the narrator via the Professional Cloning API, then render chapters through the Compose endpoint with per-segment voice and speed settings.
Outcome: You keep the original speaker's intonation across a whole book, but watch the credit math — 1 credit per character is the real budget line.
Use Cases
- Stream real-time TTS into voice agents and interactive chatbots at 170-210ms time-to-first-byte
- Generate multilingual voiceovers for e-learning and ads with automatically applied emotion
- Produce word- and character-level timestamps for auto-subtitling, karaoke highlights and lip-sync
- Clone a custom voice from a short sample via the Instant Cloning endpoint and use it in production
- Build a Professional Cloning pipeline for audiobooks where the original speaker's intonation matters
- Automate podcast and short-form narration by triggering Typecast from Zapier, Make or n8n
- Power AI tutors and AICC lines with a distinct voice per learner or caller
- Synthesize multi-speaker dialogue in a single Compose request with per-segment voice settings
Models Under the Hood
as of 2026-09-22
Limitations
- The API runs on Typecast's own ssfm-v30 model (ssfm-v21 also available), so you cannot swap in a third-party model or bring your own.
- Typecast publishes 170-210ms time-to-first-byte for streaming but not concurrency-under-load latency, so you will want to load-test with your own traffic.
- Credits are consumed on download, not on generation or playback, and burn roughly 1 credit per character — a 10-hour audiobook is not a rounding error on a Lite or Plus allowance.
- The Free plan's downloads are trial voices only, carry attribution requirements and cannot be used commercially.
- Instant and Professional Cloning slots are shared within a plan tier, and Professional Cloning is capped at 5 creations per month on Lite and 15 on Plus.
- Subscriptions bought through the iOS or Android app cannot be modified on the web and must be managed through the app store.
- Enterprise pricing, including the custom voice slot allocation and security package, is quoted per inquiry rather than published.
as of 2026-10-05
Verification history
We have re-verified Typecast AI 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 9 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Typecast AI tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0/mo
Ideal for
Solo developers or creators evaluating whether Smart Emotion and the voice library fit before spending anything.
What this tier adds
Starting tier: unlimited generation and playback, but only 3,000 lifetime download credits, trial voices only, attribution required and no commercial rights.
Basic
$5/mo billed yearly ($54/yr)
Ideal for
Individual creators monetizing YouTube, TikTok or client ad work who need commercial rights on a small budget.
What this tier adds
Adds commercial license (optional attribution), 44.1kHz downloads, access to all 600+ voices, 1 Instant Cloning slot and 30,000 credits/mo for about 35 minutes of audio.
Plus
$19/mo billed yearly ($204/yr)
Ideal for
Creators controlling pacing precisely, or anyone who needs Professional Cloning quality at a solo price.
What this tier adds
Adds speed control, the first Professional Cloning slot and 4K video export, with credits rising to 40,000/mo (~50 minutes).
Pro
$29/mo billed yearly ($312/yr)
Ideal for
Content teams whose narration needs deliberate emotional shaping rather than automatic tone alone.
What this tier adds
Unlocks Smart Emotion one-click adjustment plus custom emotion, intonation and pitch controls, 2 cloning slots total, and 75,000 credits/mo (~90 minutes).
Business
$69/mo billed yearly ($744/yr)
Ideal for
Agencies, media companies and product teams running several voice projects with multiple contributors.
What this tier adds
Adds team members and additional credit purchases, 10 cloning slots, 100GB storage and 200,000 credits/mo (~250 minutes).
Lite
$15/mo
Ideal for
Developers shipping an app or agent to production who need predictable API concurrency and cloning slots on a small monthly line item.
What this tier adds
API-specific entry tier: 200k credits/mo at $0.075 per 1k, overage at $0.09 per 1k, concurrency limit 5, and 50 shared custom voice slots with up to 2 Professional Cloning slots.
Plus (API)
$280/mo
Ideal for
Production voice agents or voiceover pipelines that routinely exceed a few million characters a month.
What this tier adds
Raises throughput to a concurrency limit of 15 and 4M credits/mo at $0.07 per 1k, with overage at $0.08 per 1k and 800 shared voice slots.
Enterprise
Custom
Ideal for
Institutions and large platforms that need API volume at scale, custom voice allocation and a named contact.
What this tier adds
Custom pricing with a dedicated account manager, customizable custom voice slots with no limits, and an enterprise API and security package.
Where the pricing makes sense
The company stage and team size where Typecast AI's pricing actually pencils out — and where peers do it cheaper.
For API work, Typecast's Lite tier at $15/mo (200k credits, 5 concurrent requests, 50 cloning slots) undercuts ElevenLabs' comparable entry tiers and sits well below Azure or Google enterprise TTS once you factor in the emotion layer you would otherwise build. The web studio runs $5/mo Basic (billed yearly at $54) up to $69/mo Business ($744/yr) for teams. Anyone past a few million characters a month should be pricing Enterprise rather than stacking Plus overage.
Setup time & first value
How long it actually takes to get something useful out of Typecast AI — broken out by persona, not the marketing-page minute.
Developers: first audio in about 10 minutes — create an API key, pick a voice ID, run the curl example or `pip install typecast` and call text_to_speech. No-code users on Zapier, Make or n8n: 20-30 minutes to wire a trigger and a Google Sheet of scripts. Web studio users on the Free plan: minutes, but downloads require upgrading to Basic at $5/mo (billed yearly at $54) before anything is
Switching to or from Typecast AI
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From XTTS v2 / self-hosted open-source TTS: swap your inference call for the text-to-speech endpoint and retire the GPU box; there is no model to convert, only request shapes.
- →From OpenAI TTS: map your voice parameter to a Typecast voice_id, then add the prompt object with emotion_type smart to get the expressiveness OpenAI leaves to you.
- →From Google or Azure TTS: port SSML-driven prosody onto Typecast's intonation, pitch and speed controls, and use Compose where you previously concatenated SSML segments.
- →From ElevenLabs: recreate custom voices through Instant or Professional Cloning from the same source samples, then repoint your SDK calls at the Typecast endpoints.
- ↗To ElevenLabs: export your source voice samples before cancelling, since cloned voices do not transfer and must be recreated on the new platform.
- ↗To a hyperscaler TTS (OpenAI, Google, Azure): you will lose the built-in emotion layer and rebuild Smart Emotion as a prompting or SSML step on your side.
- ↗To self-hosted XTTS: only worth it if you already run GPUs and your workload is high enough that per-character credits exceed infrastructure cost.
- ↗To a web-only voice editor: Typecast's API surface disappears entirely, so this is a downgrade unless you never wrote integration code.
Integrations
Resources & Guides
- Documentationtypecast.ai
Docs · Typecast AI
Full product docs from typecast.ai
- Documentationtypecast.ai
Api Reference · Typecast AI
Full product docs from typecast.ai
- Quickstarttypecast.ai
Quickstart · Typecast AI
Get up and running fast from typecast.ai
- Documentationtypecast.ai
Models · Typecast AI
Full product docs from typecast.ai
- Documentationtypecast.ai
Integrations · Typecast AI
Full product docs from typecast.ai
- Documentationtypecast.ai
Llms · Typecast AI
Full product docs from typecast.ai
- API Referencetypecast.ai
Api · Typecast AI
Methods, params, types from typecast.ai
Tutorials & Learning
YouTube returned 6 videos for “Typecast AI”, and we withheld 6: 6 could not be judged, because “Typecast AI” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Typecast AI.
Official links
Tools that pair well with Typecast AI
Common stack mates teams adopt alongside Typecast AI, with the specific reason each pairing earns its keep.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Hume AI Octave 2
Emotionally expressive text-to-speech and speech-to-speech voice AI with human-judged evaluation built in.
Noiz AI
AI voice cloning and text-to-speech with 140+ languages, emotion control, and lip-sync dubbing for creators and developers.
Featured Head-to-Head Comparisons
Typecast Ai vs Soniox
If you need a unified speech AI stack with high-accuracy STT, TTS, and translation across 60+ languages plus enterprise compliance, Soniox is the clear winner. For pure TTS with an extensive voice library, instant cloning, and a generous free tier, Typecast AI is more cost-effective and developer-friendly.
Typecast Ai vs Retell Ai
If your core need is automating phone conversations (inbound/outbound calls) with low latency and function calling for real-world tasks like booking or payments, Retell AI is the clear choice. If you need a versatile, developer-friendly TTS API with hundreds of voices, voice cloning, and multi-language support for content or app integration, Typecast AI wins. They solve different problems; the decision hinges on whether you need conversational voice agents or high-quality speech synthesis.
Typecast Ai vs Voiceitt
Choose Voiceitt if you or your users have non‑standard speech (cerebral palsy, ALS, accents) and need a dedicated dictation/captioning assistant. Choose Typecast AI if you need a developer‑friendly TTS API with 500+ voices, voice cloning, and real‑time streaming. They solve opposite problems.
Alternatives to Typecast AI
View allFish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Hume AI Octave 2
Emotionally expressive text-to-speech and speech-to-speech voice AI with human-judged evaluation built in.
Frequently Asked Questions
Categories
Best-of guides
Topics
Used Typecast AI? Help shape our editorial sentiment research.