Inworld TTS

Inworld TTS

Realtime TTS API with sub-100ms latency, instant voice cloning from 5–15 seconds of audio, and 200+ languages.

84/100Safe BetFree · from $25/moFreemium

If your product lives or dies on how fast the first syllable lands, Inworld is worth a bake-off. Realtime TTS-2 at sub-100ms plus a $25-per-1M-character entry rate that falls to $12.50 on the $1,500/mo Growth plan (or as low as $5 at enterprise scale) makes the cost math clean against ElevenLabs and Cartesia — Inworld's own June 2026 chart puts Gemini 3.1 Flash TTS at an effective ~$180/1M on a typical query. Instant cloning from 5–15 seconds and text-based voice design mean you don't need a stock library. Just budget for build-your-own voices.

Verified 5d ago · liveness 84/100 · cite: rightaichoice.com/tools/inworld-tts

Best for
  • Developers building realtime voice agents where sub-100ms TTFB decides whether the conversation feels natural
  • Teams cutting TTS spend by moving cloned or designed voices onto the $12.50–$20 per 1M character tiers
  • Multilingual consumer products needing one cloned voice to speak 200+ languages without accent carryover
  • Products with a thin stock-voice requirement that would rather clone 5–15 seconds of audio or design a voice from text
Not ideal for
  • Batch rendering of long audiobooks or podcast archives, where streaming TTS buys you nothing
  • Offline or on-device voice generation — this is a cloud API
  • Teams that want a large ready-made preset voice library off the shelf
Visit Website

IntermediateSolo developer: you can sign up free, grab an API key, clone a voice, and get first audio back from the streaming endpoint in under 30 minutes — the docs quickstart and CLI (npm install -g @inworld/cli) cover the path. Small team: budget a day to wire the Realtime API into an existing agent framework like LiveKit, Vapi, Pipecat, Mastra, or NLX. Enterprise: data residency, SLA, and DPA negotiationAPI · WebAPI availableVerified 5d ago
Pricing
Free · from $25/mo
FreemiumFree tier6 plans6 hidden costs
Learning curve
Intermediate
Solo developer: you can sign up free, grab an API key, clone a voice, and get first audio back from the streaming endpoint in under 30 minutes — the docs quickstart and CLI (npm install -g @inworld/cli) cover the path. Small team: budget a day to wire the Realtime API into an existing agent framework like LiveKit, Vapi, Pipecat, Mastra, or NLX. Enterprise: data residency, SLA, and DPA negotiation
Runs on
APIWeb
API available · 5 integrations
Who it's for
Indie voice-agent developerGame audio lead localizing a characterPlatform team consolidating voice vendors
Live sentiment
Is Inworld TTS actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Inworld TTS if you need offline or on-device synthesis, want a large ready-made preset voice library, or are batch-rendering long audio where sub-100ms streaming latency buys you nothing.

The 30-second take
Biggest gripe

Overage usage is billed at your plan's per-character rate rather than a penalty rate, so a traffic spike hits your invoice at the same $25/1M as your baseline on the free On-Demand tier — there is no cheaper bucket to

Price reality

The $25/mo Creator tier fits solo developers and small content teams; Builder at $100/mo fits small teams shipping a growing project, and Developer at $300/mo fits production applications needing 10,000 custom voices and 150 concurrent requests. Growth at $1,500/mo is where compliance and 30,000 voices live, and Enterprise goes as low as $5/1M characters with price matching — undercutting ElevenLabs and Cartesia, and far under Gemini 3.1 Flash TTS, which Inworld's own chart puts at ~$180/1M on

In short

Inworld TTS — Realtime TTS API with sub-100ms latency, instant voice cloning from 5–15 seconds of audio, and 200+ languages. Best for Developers building realtime voice agents where sub-100ms TTFB decides whether the conversation feels natural, Teams cutting TTS spend by moving cloned or designed voices onto the $12.50–$20 per 1M character tiers, Multilingual consumer products needing one cloned voice to speak 200+ languages without accent carryover. Free to start; paid plans from $25/mo.

What's new in Inworld TTS

Checked 5 days ago

Across the latest 1 update: 1 feature update.

What people actually say about Inworld TTS — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

51 mentions across 3 sources (Hacker News, YouTube, Product Hunt) · researched Aug 17, 2026.

85% positive15% critical

Average across the 3 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Sub-200ms latency for realtime streaming feels natural in voice agents
  • +Instant voice cloning from just 5-15 seconds of audio, scarily accurate
  • +Cross-lingual cloning: one voice works across 100+ languages without accent
  • +Voice steering via bracketed instructions gives granular control (tone, speed)
  • +Pricing is ~25x cheaper than ElevenLabs at enterprise scale
Recurring frustrations
  • −Lower raw quality than ElevenLabs or Minimax for premium work
  • −Free tier limited to On-Demand pricing, making heavy testing costly
  • −No built-in dubbing or media editing features like some competitors
  • −Some users report API errors when out of credits or during server load
  • −Community support is thin outside official docs
Patterns worth knowing
Unbeatable price-to-quality ratio
Seen on Hacker News, YouTube, Product Hunt
Instant voice cloning accuracy is impressive
Seen on YouTube, Product Hunt
Latency is near-invisible for realtime voice agents
Seen on Hacker News, YouTube
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • Unused characters do not roll over to next month
  • • Cloning voices on the free tier may have a watermark or require attribution
  • • Enterprise onboarding may have a minimum commitment
  • • Tokens used for voice steering count towards character usage
  • • WebSocket connections have a per-session cost at higher scale

Viability Score

84/100
Safe Bet

How well maintained and how widely used is Inworld TTS? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
85
What the vendor publishes
60

Last calculated: October 2026

How we score →

Key Features

  • Realtime streaming TTS with sub-100ms time-to-first-byte on Realtime TTS-2
  • TTS-2 Flash model tuned for speed
  • Instant voice cloning from 5–15 seconds of uploaded audio
  • Text-based voice design from natural-language descriptions of accent, age, and tone
  • Cross-lingual cloning: one voice speaks 200+ languages with no accent carryover
  • Delivery steering controls on Realtime TTS-2
  • Speaking rate and temperature controls
  • Custom pronunciation dictionary
  • Word-level timestamp alignment
  • Streaming-native WebSocket API with consistent P99 under production load
  • Multimodal realtime voice agents that can see and listen, not just talk
  • Unified Router API covering TTS, STT, and 220+ LLM models
  • Audio encodings: OGG_OPUS, WAV, and PCM
  • Voice IDs shared across TTS API, Playground, and Realtime
  • Professional voice cloning available as an add-on from the Developer tier

About Inworld TTS

FreemiumIntermediateAPI availableAPI · Web

Inworld TTS is a realtime text-to-speech API built for voice agents, games, and consumer apps where audio has to arrive before the listener notices a pause. The Realtime TTS-2 model streams audio chunks back over WebSocket as they're generated, with sub-100ms time-to-first-byte for natural conversation, and the vendor states it ranks #1 on Artificial Analysis. Beyond raw synthesis, Inworld gives you two ways to build the voice itself: instant cloning that turns 5–15 seconds of uploaded audio into a usable voice, and text-based voice design that renders a voice from a natural-language description of accent, age, and tone. Cross-lingual cloning means one voice speaks 200+ languages without accent carryover. Voice IDs are shared across the TTS API, the Playground, and the Realtime API, so a voice you design in the browser drops straight into code. The API exposes streaming and non-streaming endpoints, OGG_OPUS, WAV and PCM encoding, word-level timestamps, custom pronunciation, speaking rate and temperature controls, and delivery steering on Realtime TTS-2. A single Router API also covers 220+ LLM models and speech-to-text at $0.10–$0.15/hr, so one key and one bill cover the whole voice stack. Pricing is credit-based and metered, with volume discounts compounding at every tier: Realtime TTS-2 runs $25 per 1M characters on the free On-Demand tier down to $12.50 on Growth ($1,500/mo) and as low as $5 at enterprise scale. Inworld also announced the acquisition of Ultravox to accelerate speech-to-speech research.

Behind the Verdict

Inworld's pitch is narrow and honest: realtime voice, ranked quality, cheaper than the alternatives you're already comparing it to. The engineering that matters sits in the streaming path — audio chunks come back over a WebSocket as the model generates them, the vendor claims consistent P99 under production load, and TTFB is pegged at sub-100ms on Realtime TTS-2 (25ms on TTS-2 Flash). That's the difference between a voice agent that feels like a conversation and one that feels like a walkie-talkie. Where Inworld differs from a straight TTS vendor is the voice-construction layer. Instant cloning takes 5–15 seconds of audio and returns a usable voice; text-based voice design renders a voice from a description like "a warm, friendly female voice with a slight British accent." Cross-lingual cloning carries that voice into 200+ languages without accent carryover, which matters more for games and localization than for a single-market chatbot. Voice IDs are shared between the API, Playground, and Realtime, so there's no re-registration step when you move from prototype to production. The API surface is broad enough to be practical rather than decorative: streaming and non-streaming endpoints, OGG_OPUS/WAV/PCM output, word-level timestamps, custom pronunciation, speaking rate and temperature control, and steering controls for delivery — though steering is limited to Realtime TTS-2. The Router API folding 220+ LLM models and STT ($0.10–$0.15/hr) into one key and one bill is a real convenience for teams tired of stitching three vendors together. The honest tradeoffs. First, this is cloud-only — no offline or on-device generation. Second, Inworld pushes you toward cloned or designed voices rather than a large ready-made preset library, so if you want 50 off-the-shelf voices today, this isn't the shelf. Third, voice design and steering introduce variation, so byte-identical output across runs isn't guaranteed. Fourth, custom voice counts and API concurrency are gated by tier: 100 voices and 5 concurrent requests on free On-Demand, up to 30,000 voices and 500 concurrent requests on the $1,500/mo Growth plan. And the compliance story — ZDR, HIPAA, BAA — arrives as add-ons at the Growth tier, not as a default. On pricing cadence: monthly and yearly billing are both published, with yearly described as "2 months free," so the advertised $25/mo Creator, $100/mo Builder, $300/mo Developer and $1,500/mo Growth rates are the month-to-month figures unless you commit annually. Overage usage is charged at the same rate as your plan, which is a fairer structure than most. Bottom line: strongest fit for teams shipping realtime voice agents and multilingual consumer products; weakest fit for batch audiobook rendering where streaming gains you nothing.

Researching Inworld TTS? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Inworld TTS actually fits — and what changes day-one when you adopt it.

Indie voice-agent developer

Signs up for the free On-Demand tier, clones a voice from a 10-second sample via POST /voices/v1/voices:clone, and wires the returned voice_id into a streaming WebSocket call at api.inworld.ai/tts/v1/voice:stream with OGG_OPUS output.

Outcome: A working realtime voice agent with sub-100ms TTFB inside the 70 free TTS minutes, no card required, upgrade to Creator only when concurrency or volume outgrows it.

Game audio lead localizing a character

Designs a character voice in the Playground from a text description, then uses cross-lingual cloning so the same voice_id speaks all 15 target languages without re-recording the original actor.

Outcome: One voice identity across every locale, with the same voice_id usable in the API, Playground, and Realtime without re-registration.

Platform team consolidating voice vendors

Moves TTS, STT, and LLM calls onto the single Router API key, using STT at $0.10–$0.15/hr and 220+ LLM models on one bill instead of three vendor contracts.

Outcome: One invoice, one key, and per-tier volume discounts that compound as usage grows — TTS-2 characters fall from $25 to $12.50 by the Growth tier.

Use Cases

Models Under the Hood

Inworld Realtime TTS-2TTS-2 FlashRealtime STTinworld-tts-2

as of 2026-10-07

Limitations

  • Metered per 1M characters (TTS-2 $25/1M, TTS-2 Flash $15/1M on the base rates, discounted at higher tiers down to as low as $5/1M at enterprise scale), so high-volume synthesis carries real cost.
  • Custom voice counts and concurrency are tier-gated, from 100 custom voices on On-Demand up to 30,000 on the $1,500/mo Growth plan, and professional voice cloning is a Developer-tier add-on.
  • ZDR, HIPAA, and BAA are add-ons on the Growth plan, and EU/India data residency is Enterprise-only.
  • The per-voice character limits apply to Playground requests (40K chars on Creator).

as of 2026-10-03

Verification history

We have re-verified Inworld TTS 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 9 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Inworld TTS tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

On-Demand

$0/mo

Ideal for

Developers evaluating realtime TTS or prototyping a voice agent without committing a card — capped at 70 minutes of TTS and 100 custom voices.

What this tier adds

Free entry point: $25/1M characters for TTS-2 and $15/1M for TTS-2 Flash, 5 concurrent requests, cloning and voice design included.

Creator

$25/mo

Ideal for

Solo creators and small projects that have outgrown 70 free minutes and want a 40,000-character Playground limit per request.

What this tier adds

Adds $25/mo in credits, roughly 33% off rates ($20/1M TTS-2, $10/1M Flash), 500 custom voices, workspace sharing, and team management.

Builder

$100/mo

Ideal for

Growing projects and small teams that need more headroom than a single creator plan allows.

What this tier adds

Raises to 3,000 custom voices and 50 concurrent requests, with rates at $17.50/1M TTS-2 and $9/1M Flash — up to 40% off.

Developer

$300/mo

Ideal for

Production applications with real traffic that need 10,000 custom voices and 150 concurrent requests.

What this tier adds

Unlocks professional voice cloning as an add-on, priority email support, and $15/1M TTS-2 rates — up to 47% off.

Growth

$1,500/mo

Ideal for

Large deployments with compliance requirements — the first tier where ZDR, HIPAA, and BAA are available as add-ons.

What this tier adds

30,000 custom voices, 500 concurrent requests, one professional voice clone included, and $12.50/1M TTS-2 rates — up to 53% off.

Enterprise

Custom

Ideal for

Organizations needing custom limits, SLA and DPA coverage, EU or India data residency, and a dedicated account manager.

What this tier adds

Custom terms with TTS-2 as low as $5/1M characters and TTS-2 Flash sub-$5/1M, plus price matching against competing vendors.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Overage usage is billed at your plan's per-character rate rather than a penalty rate, so a traffic spike hits your invoice at the same $25/1M as your baseline on the free On-Demand tier — there is no cheaper bucket to
  • Concurrent requests are capped at 5 on the free On-Demand tier, 10 on Creator ($25/mo), 50 on Builder ($100/mo), and 150 on Developer ($300/mo), so a burst of simultaneous users can queued-block well before your
  • Custom voices are capped by tier — 100 on free, 500 on Creator, 3,000 on Builder, 10,000 on Developer — and building a large voice catalog means paying for a higher plan even if your synthesis volume stays flat.
  • Professional voice cloning is an add-on that only unlocks from the Developer tier ($300/mo) up, so a small team that needs a single pro clone can't get it on the $25 or $100 plans.
  • ZDR, HIPAA, and BAA are add-ons locked to the $1,500/mo Growth tier, so regulated workloads can't stay on cheaper plans regardless of usage.
  • The monthly prices shown ($25/$100/$300/$1,500 per month) are the month-to-month rates — the yearly option is advertised as '2 months free', so the effective annual rate is lower, but you pay it up front.

Where the pricing makes sense

The company stage and team size where Inworld TTS's pricing actually pencils out — and where peers do it cheaper.

The $25/mo Creator tier fits solo developers and small content teams; Builder at $100/mo fits small teams shipping a growing project, and Developer at $300/mo fits production applications needing 10,000 custom voices and 150 concurrent requests. Growth at $1,500/mo is where compliance and 30,000 voices live, and Enterprise goes as low as $5/1M characters with price matching — undercutting ElevenLabs and Cartesia, and far under Gemini 3.1 Flash TTS, which Inworld's own chart puts at ~$180/1M on

Setup time & first value

How long it actually takes to get something useful out of Inworld TTS — broken out by persona, not the marketing-page minute.

Solo developer: you can sign up free, grab an API key, clone a voice, and get first audio back from the streaming endpoint in under 30 minutes — the docs quickstart and CLI (npm install -g @inworld/cli) cover the path. Small team: budget a day to wire the Realtime API into an existing agent framework like LiveKit, Vapi, Pipecat, Mastra, or NLX. Enterprise: data residency, SLA, and DPA negotiation

Switching to or from Inworld TTS

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From ElevenLabs: map existing voice IDs to Inworld clones (5–15 seconds of sample audio each), then swap the endpoint to api.inworld.ai/tts/v1/voice:stream and set model_id to inworld-tts-2.
  • →From Cartesia: rebuild voices with instant cloning or text-based voice design, then point your streaming client at the Inworld WebSocket endpoint and pick OGG_OPUS or PCM output.
  • →From Google Gemini 3.1 Flash TTS: replace token-based audio-output billing with per-character metering — Inworld prices at $25/1M characters on On-Demand versus roughly $180/1M effective on a typical Gemini query.
  • →From a self-hosted open-source TTS: keep your text pipeline, replace the synthesis call with Inworld's streaming endpoint and offload latency and scaling to the API.
Migrating out
  • ↗To ElevenLabs: if you need a much larger ready-made preset voice library, re-clone or re-select voices and swap the synthesis endpoint.
  • ↗To Cartesia: export your voice samples and rebuild them on Cartesia's cloning flow if you need a different latency or pricing profile.
  • ↗To a self-hosted TTS model: run your own inference if offline or on-device synthesis is a requirement, accepting the infrastructure and quality tradeoff.
  • ↗To Google Gemini 3.1 Flash TTS: move to token-based audio-output billing if you are already committed to the Gemini stack, at a materially higher effective per-1M-character rate.

Integrations

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Inworld TTS”, and we withheld 3: 3 did not mention Inworld TTS. Showing the 3 we can prove are about Inworld TTS.

Tools that pair well with Inworld TTS

Common stack mates teams adopt alongside Inworld TTS, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Inworld TTS

View all
Fish Audio

Fish Audio

Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.

FreemiumTry
Speechify Studio - AI Voice Generator

Speechify Studio - AI Voice Generator

AI voice generator with 1,000+ lifelike voices in 60+ languages, plus dubbing, cloning, avatars, and captions

FreemiumTry
Noiz AI

Noiz AI

AI voice cloning and text-to-speech with 140+ languages, emotion control, and lip-sync dubbing for creators and developers.

PaidTry

Frequently Asked Questions

Used Inworld TTS? Help shape our editorial sentiment research.