Typecast AI

Typecast AI

Expressive text-to-speech API with Smart Emotion, 600+ voices in 35+ languages, and 170-210ms streaming for real-time agents.

79/100Safe BetFree · from $5/mo billed yearly ($54/yr)Freemium

If your agent needs to sound like it means what it says, Typecast's Smart Emotion plus 170-210ms streaming is the combination most competitors make you assemble by hand, and the September 2026 remove_silence_ms parameter shows the team is still shipping developer-facing polish rather than voice-count marketing. Against OpenAI's or Azure's TTS you get more explicit emotion control and a lower published unit cost; against ElevenLabs you give up some brand recognition and the on-prem option, and you cannot buy a plan that follows you across iOS, Android and web. Evaluation is genuinely free, but commercial rights only start at $5/mo billed annually ($54/yr) on Basic, and long-form audio still

Verified 3d ago · liveness 79/100 · cite: rightaichoice.com/tools/typecast-ai

Best for
  • Developers building voice agents that need sub-200ms streaming TTS
  • Content teams automating YouTube, TikTok and ad voiceovers without a studio
  • Audiobook and podcast producers who need expressive narration with emotion control
  • Product teams adding multilingual narration to an existing app
Not ideal for
  • Teams that require on-premise or air-gapped deployment
  • Long-form batch rendering where per-character credit cost dominates the budget
  • Projects needing unlimited commercial output on a zero budget
Visit Website

IntermediateDevelopers: first audio in about 10 minutes — create an API key, pick a voice ID, run the curl example or `pip install typecast` and call text_to_speech. No-code users on Zapier, Make or n8n: 20-30 minutes to wire a trigger and a Google Sheet of scripts. Web studio users on the Free plan: minutes, but downloads require upgrading to Basic at $5/mo (billed yearly at $54) before anything isWeb · API · CLIAPI availableVerified 3d ago
Pricing
Free · from $5/mo billed yearly ($54/yr)
FreemiumFree tier8 plans6 hidden costs
Learning curve
Intermediate
Developers: first audio in about 10 minutes — create an API key, pick a voice ID, run the curl example or `pip install typecast` and call text_to_speech. No-code users on Zapier, Make or n8n: 20-30 minutes to wire a trigger and a Google Sheet of scripts. Web studio users on the Free plan: minutes, but downloads require upgrading to Basic at $5/mo (billed yearly at $54) before anything is
Runs on
WebAPICLI
API available · 10 integrations
Who it's for
Backend developer building a voice agentContent producer running short-form videoAudiobook producer replacing a studio session
Live sentiment
Is Typecast AI actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Typecast if you need on-premise or air-gapped deployment, or if you are optimizing purely on per-character cost at very large batch volume.

The 30-second take
Biggest gripe

Credits are only consumed when you download audio — generation and playback are free — so budgeting by trial usage will understate your real bill.

Price reality

For API work, Typecast's Lite tier at $15/mo (200k credits, 5 concurrent requests, 50 cloning slots) undercuts ElevenLabs' comparable entry tiers and sits well below Azure or Google enterprise TTS once you factor in the emotion layer you would otherwise build. The web studio runs $5/mo Basic (billed yearly at $54) up to $69/mo Business ($744/yr) for teams. Anyone past a few million characters a month should be pricing Enterprise rather than stacking Plus overage.

In short

Typecast AI — Expressive text-to-speech API with Smart Emotion, 600+ voices in 35+ languages, and 170-210ms streaming for real-time agents. Best for Developers building voice agents that need sub-200ms streaming TTS, Content teams automating YouTube, TikTok and ad voiceovers without a studio, Audiobook and podcast producers who need expressive narration with emotion control. Free to start; paid plans from $5/mo.

What's new in Typecast AI

Checked 3 days ago

Across the latest 5 updates: 3 feature updates, 1 launch and 1 changelog entry.

What people actually say about Typecast AI — is it worth it?

We scanned public community sources for Typecast AI on Sep 24, 2026 and could not establish that the discussion we found is about this tool rather than something else sharing its name. Only 1 of the posts we fetched could be positively tied to Typecast AI. Rather than publish a sentiment score built on the wrong subject, we publish nothing here and re-run the scan.

Viability Score

79/100
Safe Bet

How well maintained and how widely used is Typecast AI? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
35
What the vendor publishes
60

Last calculated: October 2026

How we score →

Key Features

  • Text-to-speech API (POST /v1/text-to-speech) returning WAV or MP3
  • Streaming TTS endpoint with 170-210ms time-to-first-byte for real-time agents
  • Timestamp TTS with word- and character-level alignment for subtitles and karaoke
  • Compose TTS endpoint synthesizing multiple segments in one request with per-segment settings
  • Smart Emotion applies tone automatically from surrounding context (previous_text / next_text)
  • 7 manual emotion presets plus custom emotion, intonation, pitch and speed controls
  • remove_silence_ms parameter (0-1000ms) to trim detected silence without cutting explicit pauses
  • Instant voice cloning from a short WAV or MP3 sample
  • Professional Voice Cloning API (async) for high-accuracy intonation and tone matching
  • Voice recommendations endpoint matching voices to a natural-language description
  • 600+ voices across ages, tones and personalities, 35+ languages with native-level naturalness in 6
  • Official SDKs for Python, JavaScript/TypeScript, Go, Rust, C#, Java, Kotlin, C, Swift, Zig, PHP, Dart and Ruby
  • Cast CLI (v1.0.9) for command-line generation and agent workflows
  • No-code integrations via Zapier, Make, n8n and Google Sheets
  • Agent integrations including MCP server, OpenClaw, Claude Skills and Pipecat

About Typecast AI

FreemiumIntermediateAPI availableWeb · API · CLI

Typecast is a cloud text-to-speech API from Neosapience, aimed at developers building conversational agents, content teams automating voiceovers, and product teams adding narration to apps. You send text to a REST endpoint and get back WAV or MP3 audio, with official SDKs covering Python, JavaScript, Go, Rust, C#, Java, Kotlin, C, Swift, Zig, PHP, Dart and Ruby, a Cast CLI (v1.0.9 as of the August 2026 release), and no-code hooks into Zapier, Make, n8n and Google Sheets if you would rather not write glue code. The pitch rests on expressiveness rather than raw voice count, though the count is real: 600+ voices spanning 35+ languages, with native-level naturalness claimed in 6, English included. Smart Emotion reads surrounding context — you can pass previous_text and next_text — and applies tone automatically; 7 manual emotion presets plus custom emotion, intonation, pitch and speed controls let you override when you want a specific read. Voice cloning comes in two flavors: Instant Cloning from a short sample, and Professional Cloning (launched August 2026 as an async API endpoint) that chases the original speaker's intonation and tone. The underlying model is Typecast's own SSFM line, with ssfm-v30 adding context-based emotion control and ssfm-v21 still available. For real-time work there is a Streaming TTS endpoint reporting 170-210ms time-to-first-byte depending on model, which is the number that matters if you are wiring voice into a live agent loop rather than batch-rendering files. A Timestamp endpoint returns word- and character-level alignments for subtitles, karaoke and lip-sync, and the July 2026 'Compose' endpoint synthesizes multiple segments in one request with per-segment voice and speech settings. The September 2026 release added a remove_silence_ms parameter (0-1000ms) across TTS, streaming TTS, timestamped TTS and Compose, now propagated to the official SDKs, Cast CLI, n8n, Zapier, Pipecat and the hosted Typecast API MCP server. Positioning is straightforward: cheaper and more curated than a global TTS API, more turnkey than assembling XTTS v2 yourself on your own GPUs. If you need on-prem or air-gapped deployment, or you are optimizing purely on per-character cost at very large volume, look elsewhere first.

Behind the Verdict

Typecast's real differentiation is not the 600-voice library — it is that emotion is a first-class API parameter rather than something you bolt on. Smart Emotion inspects the surrounding text (you literally pass previous_text and next_text) and picks the tone, and when you need a specific read you drop to one of 7 manual presets or set custom emotion, intonation, pitch and speed. For conversational agents this matters more than voice variety, because a chatbot that reads every line flatly sounds broken no matter how good the timbre is. The developer surface is unusually broad for a company this size: REST endpoints for full TTS, streaming TTS, timestamped TTS, Compose (multi-segment in one request), instant cloning, professional cloning, and voice recommendations from a natural-language description (GET /v1/voices/recommendations, July 2026). Official SDKs span Python, JS/TS, Go, Rust, C#, Java, Kotlin, C, Swift, Zig, PHP, Dart and Ruby — a list most TTS vendors do not come close to — plus a Cast CLI and a docs site that ships an llms.txt so coding agents can read it directly. MCP, Claude Skills, OpenClaw and Pipecat integrations put it in the agent stack rather than beside it. Where it gets thinner: the model is ssfm-v30 (or ssfm-v21), and Typecast does not publish concurrency-under-load latency figures, only the 170-210ms TTFB range. There is no on-prem or air-gapped deployment. The credit model is character-based, so a 10-hour audiobook at 1 credit per character is a real budget line, and Enterprise is custom-quoted with no published number. Longer-term it is a Neosapience product, so you are betting on a specialist rather than a hyperscaler whose TTS ships as one line item among hundreds. Where it fits: voice agents, AI tutors, IVR/AICC, subtitled video pipelines, and any team that wants expressive narration without hiring a studio or babysitting a GPU. Where it does not: teams with a strict data-residency requirement, anyone whose only metric is dollars per million characters at massive scale, and buyers who expect one subscription to cover web plus both mobile apps.

Researching Typecast AI? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Typecast AI actually fits — and what changes day-one when you adopt it.

Backend developer building a voice agent

You point the Streaming TTS endpoint at your LLM's output stream, set emotion_type to smart, and pass previous_text so the tone carries across turns.

Outcome: Your agent starts speaking within 170-210ms of the first chunk and sounds like it means what it says, without a separate emotion-classification step.

Content producer running short-form video

You trigger Typecast from Zapier or Make with a script row from Google Sheets, then pull the timestamped TTS output into your captioning step.

Outcome: Voiceovers and word-level caption timings arrive in one pipeline, and remove_silence_ms trims the dead air without cutting the deliberate pauses.

Audiobook producer replacing a studio session

You clone the narrator via the Professional Cloning API, then render chapters through the Compose endpoint with per-segment voice and speed settings.

Outcome: You keep the original speaker's intonation across a whole book, but watch the credit math — 1 credit per character is the real budget line.

Use Cases

Models Under the Hood

ssfm-v30 (SSFM 3.0)ssfm-v21

as of 2026-09-22

Limitations

  • The API runs on Typecast's own ssfm-v30 model (ssfm-v21 also available), so you cannot swap in a third-party model or bring your own.
  • Typecast publishes 170-210ms time-to-first-byte for streaming but not concurrency-under-load latency, so you will want to load-test with your own traffic.
  • Credits are consumed on download, not on generation or playback, and burn roughly 1 credit per character — a 10-hour audiobook is not a rounding error on a Lite or Plus allowance.
  • The Free plan's downloads are trial voices only, carry attribution requirements and cannot be used commercially.
  • Instant and Professional Cloning slots are shared within a plan tier, and Professional Cloning is capped at 5 creations per month on Lite and 15 on Plus.
  • Subscriptions bought through the iOS or Android app cannot be modified on the web and must be managed through the app store.
  • Enterprise pricing, including the custom voice slot allocation and security package, is quoted per inquiry rather than published.

as of 2026-10-05

Verification history

We have re-verified Typecast AI 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-checked, vendor evidence unchanged
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 9 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Typecast AI tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free

$0/mo

Ideal for

Solo developers or creators evaluating whether Smart Emotion and the voice library fit before spending anything.

What this tier adds

Starting tier: unlimited generation and playback, but only 3,000 lifetime download credits, trial voices only, attribution required and no commercial rights.

Basic

$5/mo billed yearly ($54/yr)

Ideal for

Individual creators monetizing YouTube, TikTok or client ad work who need commercial rights on a small budget.

What this tier adds

Adds commercial license (optional attribution), 44.1kHz downloads, access to all 600+ voices, 1 Instant Cloning slot and 30,000 credits/mo for about 35 minutes of audio.

Plus

$19/mo billed yearly ($204/yr)

Ideal for

Creators controlling pacing precisely, or anyone who needs Professional Cloning quality at a solo price.

What this tier adds

Adds speed control, the first Professional Cloning slot and 4K video export, with credits rising to 40,000/mo (~50 minutes).

Pro

$29/mo billed yearly ($312/yr)

Ideal for

Content teams whose narration needs deliberate emotional shaping rather than automatic tone alone.

What this tier adds

Unlocks Smart Emotion one-click adjustment plus custom emotion, intonation and pitch controls, 2 cloning slots total, and 75,000 credits/mo (~90 minutes).

Business

$69/mo billed yearly ($744/yr)

Ideal for

Agencies, media companies and product teams running several voice projects with multiple contributors.

What this tier adds

Adds team members and additional credit purchases, 10 cloning slots, 100GB storage and 200,000 credits/mo (~250 minutes).

Lite

$15/mo

Ideal for

Developers shipping an app or agent to production who need predictable API concurrency and cloning slots on a small monthly line item.

What this tier adds

API-specific entry tier: 200k credits/mo at $0.075 per 1k, overage at $0.09 per 1k, concurrency limit 5, and 50 shared custom voice slots with up to 2 Professional Cloning slots.

Plus (API)

$280/mo

Ideal for

Production voice agents or voiceover pipelines that routinely exceed a few million characters a month.

What this tier adds

Raises throughput to a concurrency limit of 15 and 4M credits/mo at $0.07 per 1k, with overage at $0.08 per 1k and 800 shared voice slots.

Enterprise

Custom

Ideal for

Institutions and large platforms that need API volume at scale, custom voice allocation and a named contact.

What this tier adds

Custom pricing with a dedicated account manager, customizable custom voice slots with no limits, and an enterprise API and security package.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Credits are only consumed when you download audio — generation and playback are free — so budgeting by trial usage will understate your real bill.
  • Going past your monthly allowance is pay-as-you-go at $0.09 per 1k credits on Lite and $0.08 per 1k on Plus, which compounds quickly on long-form narration.
  • The Free plan's 3,000 download credits are lifetime, not monthly, and every download is trial-voice-only with attribution required and no commercial rights.
  • Commercial use requires at least the Basic tier at $5/mo, and that rate is the yearly price ($54 billed annually) — month-to-month costs more.
  • Instant and Professional Cloning draw from the same slot pool, and Professional Cloning creations are capped at 5 per month on Lite and 15 on Plus.
  • Plans bought through the iOS or Android app cannot be changed or cancelled on the web — you have to go back through the app store.

Where the pricing makes sense

The company stage and team size where Typecast AI's pricing actually pencils out — and where peers do it cheaper.

For API work, Typecast's Lite tier at $15/mo (200k credits, 5 concurrent requests, 50 cloning slots) undercuts ElevenLabs' comparable entry tiers and sits well below Azure or Google enterprise TTS once you factor in the emotion layer you would otherwise build. The web studio runs $5/mo Basic (billed yearly at $54) up to $69/mo Business ($744/yr) for teams. Anyone past a few million characters a month should be pricing Enterprise rather than stacking Plus overage.

Setup time & first value

How long it actually takes to get something useful out of Typecast AI — broken out by persona, not the marketing-page minute.

Developers: first audio in about 10 minutes — create an API key, pick a voice ID, run the curl example or `pip install typecast` and call text_to_speech. No-code users on Zapier, Make or n8n: 20-30 minutes to wire a trigger and a Google Sheet of scripts. Web studio users on the Free plan: minutes, but downloads require upgrading to Basic at $5/mo (billed yearly at $54) before anything is

Switching to or from Typecast AI

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From XTTS v2 / self-hosted open-source TTS: swap your inference call for the text-to-speech endpoint and retire the GPU box; there is no model to convert, only request shapes.
  • →From OpenAI TTS: map your voice parameter to a Typecast voice_id, then add the prompt object with emotion_type smart to get the expressiveness OpenAI leaves to you.
  • →From Google or Azure TTS: port SSML-driven prosody onto Typecast's intonation, pitch and speed controls, and use Compose where you previously concatenated SSML segments.
  • →From ElevenLabs: recreate custom voices through Instant or Professional Cloning from the same source samples, then repoint your SDK calls at the Typecast endpoints.
Migrating out
  • ↗To ElevenLabs: export your source voice samples before cancelling, since cloned voices do not transfer and must be recreated on the new platform.
  • ↗To a hyperscaler TTS (OpenAI, Google, Azure): you will lose the built-in emotion layer and rebuild Smart Emotion as a prompting or SSML step on your side.
  • ↗To self-hosted XTTS: only worth it if you already run GPUs and your workload is high enough that per-character credits exceed infrastructure cost.
  • ↗To a web-only voice editor: Typecast's API surface disappears entirely, so this is a downgrade unless you never wrote integration code.

Integrations

ZapierMaken8nGoogle SheetsOpenClawClaude SkillsMCPPipecatLlamaIndexPostman

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Typecast AI”, and we withheld 6: 6 could not be judged, because “Typecast AI” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Typecast AI.

Tools that pair well with Typecast AI

Common stack mates teams adopt alongside Typecast AI, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Typecast AI

View all
Fish Audio

Fish Audio

Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.

FreemiumTry
Hume AI Octave 2

Hume AI Octave 2

Emotionally expressive text-to-speech and speech-to-speech voice AI with human-judged evaluation built in.

FreemiumTry
Noiz AI

Noiz AI

AI voice cloning and text-to-speech with 140+ languages, emotion control, and lip-sync dubbing for creators and developers.

PaidTry

Frequently Asked Questions

Used Typecast AI? Help shape our editorial sentiment research.