Voicemaker
Voicemaker turns text into 140-language AI voiceovers with music generation, cloning, and one credit balance.
Voicemaker's pitch is volume and range at a low per-credit cost, and the August 2026 lineup — SuperTTS, Pulse 1.0, SuperSeed 1.0, Mercury 1.0 — backs it up better than the old ProV2-only story did. Mercury 1.0 generating BGM and song scores from the same account you do voiceovers from is a genuinely unusual combination at $5–$20/mo. It won't win a blind listening test against ElevenLabs, and it shouldn't try to. Buy it when you need many languages, many minutes, and one bill. Skip it if your whole brief is the single most human-sounding voice on the market, or if you need offline/on-premise TTS without a custom enterprise deal.
Verified 14d ago · liveness 68/100 · cite: rightaichoice.com/tools/voicemaker
- Creators dubbing video, podcast and ad content into multiple languages from one credit balance
- E-learning and audiobook publishers who need long-form narration plus SRT subtitle files
- Developers wiring TTS, STT or Speech-to-Speech into a product via REST API
- Teams producing IVR and chatbot audio that must work across many languages and accents
- Buyers whose entire brief is the single most human-sounding voice on the market
- Anyone needing offline, desktop, or on-premise TTS without a custom enterprise deal
- High-volume Speech-to-Speech workflows — that runs 100 credits per second and drains allowances fast
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Voicemaker if the voice itself is the deliverable and you need the most human-sounding audio at any price, or if you need offline/on-premise TTS without a custom enterprise contract.
Speech-to-Speech burns 100 credits per second, so a five-minute redub costs 30,000 credits — a meaningful bite out of a Creator plan's 500,000 monthly allowance.
Voicemaker's $5–$30/mo range undercuts ElevenLabs and PlayHT on per-character cost while bundling music, subtitle export and STT into the same balance — so solo creators and small teams get the most value on Starter ($5), Creator ($10) or Pro ($20). Audiobook and podcast publishers should take the $25/year dedicated plan (1M credits, 100,000-char conversions). Teams and Business sit at $30–$50/mo with shared workspaces and admin controls. Business/Enterprise is custom-priced with SSO and volume
In short
Voicemaker — Voicemaker turns text into 140-language AI voiceovers with music generation, cloning, and one credit balance. Best for Creators dubbing video, podcast and ad content into multiple languages from one credit balance, E-learning and audiobook publishers who need long-form narration plus SRT subtitle files, Developers wiring TTS, STT or Speech-to-Speech into a product via REST API. Free to start; paid plans from $5/mo.
What's new in Voicemaker
Checked todayAcross the latest 2 updates: 1 feature update and 1 changelog entry.
Voicemaker adds 70 voices, credit rollover and redesigned Teams/Business admin console
50 new ProPlus voices and 20 FlashX voices. Unused credits roll over up to 1 month on Pro, 3 months on Teams and Business.
Voicemaker v1.9.2 adds SuperTTS all-in-one audio model and three new generation modes
SuperTTS generates voice, ambience and effects from one prompt. New modes Pulse 1.0, SuperSeed 1.0 and Mercury 1.0 cover 18-70+ languages each.
Viability Score
How well maintained and how widely used is Voicemaker? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Text-to-speech in 140 languages with regional accents
- SuperTTS: voice, ambience and effects generated from one prompt
- Pulse 1.0 consistency model — same prompt, same audio (70+ languages)
- SuperSeed 1.0 model for dialogue-heavy multi-speaker scenes (18+ languages)
- Mercury 1.0 model for songs, singing and background music (70+ languages)
- Multi-Speaker Mode assigns different voices to specific dialogue lines
- Studio Grade Stitching for long-form narration (auto-enables past 1,000 chars)
- Speech-to-Speech voice conversion and redubbing
- Speech-to-Text in 90+ languages with SRT and TXT subtitle export
- Voice cloning slots on Pro and above; extra slots at $2/year
- VoxFX with 100+ voice effects, free for unlimited conversions
- VoxStudio suite with Voice Enhancer, Voice Isolator and Music Sense
- SSML support, Pronunciation Editor, pause/pitch/speed/volume sliders
- REST API for TTS, STT and STS with pay-as-you-go credits
- Audio output in MP3, WAV, OGG, AAC, OPUS, ULAW up to 48kHz and 8kHz telephony
About Voicemaker
Voicemaker is a cloud text-to-speech platform for people who ship a lot of audio and don't want to pay premium per-seat prices for it. You paste text or upload a PDF/DOC/TXT, pick from a 1,000+ voice library across 140 languages, and export MP3, WAV (16-bit PCM 48kHz), OGG, AAC, OPUS, ULAW, or 8kHz telephony audio. One currency — credits — covers every tool and model, so a voiceover, a transcription, and a speech-to-speech redub all draw from the same balance. The 2026 model lineup goes past straight narration. Pulse 1.0 is the consistency play: the same prompt produces the same audio every time across 70+ languages, with SFX and BGM placed precisely in the prompt — the pick for ads, podcast intros, and episodic content. SuperSeed 1.0 blends Pulse-level consistency with Mercury-level creativity for multi-speaker, dialogue-heavy scenes in 18+ languages. Mercury 1.0 targets songs, singing, and background scores with a fresh take on each run in 70+ languages. SuperTTS generates voice, ambience, and effects from a single prompt for cinematic scenes. Long-form narration gets Studio Grade Stitching, which auto-enables past 1,000 characters. Beyond TTS the toolkit includes Speech-to-Speech redubbing, Speech-to-Text in 90+ languages with SRT/TXT subtitle export, VoxStudio for cleanup with Voice Enhancer and Voice Isolator, VoxFX with 100+ effects, an SSML pronunciation editor, Multi-Speaker Mode for dialogue, and voice cloning slots. Developers get REST APIs for TTS, STT, and STS on pay-as-you-go credits. The positioning is straightforward: Voicemaker competes on breadth of models and cost per character, not on the absolute realism crown. If your priority is the single most human-sounding voice at any price, ElevenLabs and PlayHT are the usual comparisons. If you need 140 languages, music generation, subtitle export, and cloning under one credit balance, the math here is hard to argue with.
Behind the Verdict
The strongest thing about Voicemaker in late 2026 isn't any single voice — it's that you can produce a narrated video, its subtitles, its background score, and a redubbed version in another language without leaving the account or opening a second vendor relationship. Credits work across all tools and models, so a SuperTTS cinematic scene, a Pulse 1.0 ad read, a Speech-to-Text pass for subtitles, and a Mercury 1.0 BGM bed all come out of one balance. For a small studio or a solo creator dubbing YouTube content, that consolidation is the product. The model lineup is now genuinely differentiated rather than a single engine with new labels. Pulse 1.0 exists because production teams kept getting burned by non-deterministic output — same prompt, same audio, every time, 70+ languages, with SFX and BGM placement you control in the prompt. SuperSeed 1.0 trades a little of that lockstep for expressive variation in dialogue-heavy scenes across 18+ languages. Mercury 1.0 is the odd one out: it makes music and soundscapes, not narration, and it's the reason a podcast producer can score an episode here instead of buying a second subscription. SuperTTS bundles voice, ambience, and effects into one prompt for cinematic ads. The credit economics are the other half of the pitch, and they cut both ways. Default voices and FlashX cost 1 credit per character; ProV1 and ProV2 cost 2; ProPlus Turbo costs 2; ProPlus Expressive and High-Res cost 4. Speech-to-Speech burns 100 credits per second — a five-minute redub is 30,000 credits, which eats a Creator plan's 500,000 monthly allowance fast. Speech-to-Text is 10 credits per second. So a plan that advertises '~9 hours of audio' really means nine hours of Default-voice narration, not nine hours of redubbing. If your workflow is mostly voice conversion, do the arithmetic before you pick a tier. Feature gating is where the low headline price gets less low. Pronunciation Editor, Voice Profile, VoxStudio, VoxFX, Projects, SSML, and SRT export are all locked to paid plans starting at Creator ($10/mo). Voice cloning, Speech-to-Speech, and Speech-to-Text don't appear until Pro ($20/mo). 2FA is a Creator-and-above feature. Credit rollover only exists on Pro (1 billing cycle) and Teams/Business (up to 3 months) — Starter and Creator users lose unused credits every month. None of that is hidden, but it means the $5 tier is closer to a trial than a working plan. The realism question still has an honest answer: if the voice itself is the deliverable — a premium brand film, an audiobook where a listener will notice every artifact — ElevenLabs and PlayHT remain the safer picks. Voicemaker's own ecosystem acknowledges this by shipping cleanup tools (Voice Enhancer, Voice Isolator) and Studio Grade Stitching rather than claiming the realism crown outright. Where it does win is breadth: 140 languages, 1,000+ voices, six output formats, an SSML editor, and REST APIs for TTS/STT/STS at pay-as-you-go rates that undercut the major clouds
Researching Voicemaker? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Voicemaker actually fits — and what changes day-one when you adopt it.
You write a 1,200-character script in English, run it through Pulse 1.0 for a repeatable intro voice, generate a Mercury 1.0 BGM bed, then push the whole thing through Speech-to-Text to export matching .SRT subtitles.
Outcome: A narrated video with a scored intro and synced subtitles ships from one account in an afternoon, with all four tools drawing from a single Creator-plan credit balance.
You upload a 40-page course PDF, let Studio Grade Stitching auto-enable for the long lessons, assign distinct ProPlus voices to a two-person dialogue in Multi-Speaker Mode, and export WAV 48kHz plus SRT per lesson.
Outcome: A full module ships with consistent narration and downloadable subtitles, avoiding eleven separate TTS sessions and a separate transcription vendor.
You call the REST API for TTS in a mobile app, add STS to redub user recordings, and pull STT for on-screen captions — all on pay-as-you-go credits.
Outcome: Voice features go into production with commercial rights included, documented code samples, and usage scaling without a negotiated seat contract.
Use Cases
- Generate multilingual voiceovers for YouTube videos with accent and emotion control.
- Create interactive IVR systems for customer support using customisable 8kHz telephony voice prompts.
- Produce audiobooks with expressive narration using Studio Grade Stitching for long chapters.
- Build conversational AI agents with Multi-Speaker Mode for dialogue between different characters.
- Develop e-learning courses with consistent narration across 140 languages, plus SRT subtitles.
- Score ads, podcast intros and cinematic scenes with Mercury 1.0 or SuperTTS from the same credit balance.
Models Under the Hood
as of 2026-09-29
Limitations
- Credits are consumed per character or per second depending on the tool and voice model.
- Default voices (AI1, AI2, AI3, AI4) and FlashX cost 1 credit per character; ProPlus Turbo and ProV2 cost 2; ProPlus Expressive and High-Res cost 4; Speech-to-Speech runs 100 credits per second and Speech-to-Text 10 credit per second.
- Pause settings are supported only for Default voices (AI1, AI2, AI3, AI4) and Pro voices (ProPlus – High-Res & Turbo VoiceModel, and ProV1), and Pronunciation Editor, Voice Profile, subtitles and voice cloning require a paid plan.
- Credit rollover is limited: up to 1 month on Pro and up to 3 months on Teams and Business.
as of 2026-09-15
Verification history
We have re-verified Voicemaker 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Voicemaker tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0/forever
Ideal for
Someone testing whether Voicemaker's output quality fits their project before committing to a paid plan.
What this tier adds
Free entry point: 250 characters per conversion, limited generations per day, 120 languages, MP3/OGG 128 kbps, personal use only.
Starter
$5/mo
Ideal for
Hobbyist creators producing occasional voiceovers who don't yet need clone, STT or advanced editing tools.
What this tier adds
Adds 200,000 credits/month (~4 hours), 3,000-character conversions, 500+ voice library, 140 languages, and commercial usage rights for $5/mo.
Creator
$10/mo
Ideal for
Working creators shipping premium content for global audiences who need editing and subtitle tools.
What this tier adds
Adds VoxStudio, VoxFX, Projects, SSML, Pronunciation Editor, SRT export, cloud storage and 2FA for $10/mo, plus 500,000 credits (~9 hours).
Pro
$20/mo
Ideal for
Professionals scaling production who need voice cloning, redubbing and transcription in the same account.
What this tier adds
Unlocks voice cloning, Speech-to-Speech and Speech-to-Text, 1M credits/month (~18 hours), 10,000-char conversions, and 1-cycle credit rollover for $20/mo.
Teams
$30/mo
Ideal for
Small teams collaborating on shared AI audio projects with a need for usage governance.
What this tier adds
Adds a shared workspace, admin console, workspace-level file sharing, 3-month credit rollover, and $20/year extra seats on top of the Pro feature set for $30/mo.
Audiobook & Podcast Creation
$25/year
Ideal for
Publishers and podcasters producing long-form narration who prefer an annual commitment.
What this tier adds
Annual dedicated plan at $25/year with 1M credits, 100,000-character conversions, 1,000+ voices, SSML, 10 GB storage and YouTube-ready exports.
Developer API
Pay-as-you-go
Ideal for
Product teams embedding TTS, STT or Speech-to-Speech into an app and paying by usage.
What this tier adds
Pay-as-you-go REST access to the full voice library including Expressive and cloning, full docs and code samples, and commercial use included.
Business / Enterprise
Custom
Ideal for
Larger organisations needing SSO, custom volumes and negotiated contract terms.
What this tier adds
Adds enterprise admin and SSO login, custom credit volumes and seats, volume top-ups at $20 per 1M credits, and extra clone slots at $2/year.
Where the pricing makes sense
The company stage and team size where Voicemaker's pricing actually pencils out — and where peers do it cheaper.
Voicemaker's $5–$30/mo range undercuts ElevenLabs and PlayHT on per-character cost while bundling music, subtitle export and STT into the same balance — so solo creators and small teams get the most value on Starter ($5), Creator ($10) or Pro ($20). Audiobook and podcast publishers should take the $25/year dedicated plan (1M credits, 100,000-char conversions). Teams and Business sit at $30–$50/mo with shared workspaces and admin controls. Business/Enterprise is custom-priced with SSO and volume
Setup time & first value
How long it actually takes to get something useful out of Voicemaker — broken out by persona, not the marketing-page minute.
Individual creators: about 10 minutes to register, pick a Default voice, paste text and download an MP3 at the free tier. Paid-plan workflows with cloning, VoxStudio or projects: 30–60 minutes to set up a voice and learn the credit rates. Developers: budget 1–2 hours to read the REST docs, grab an API key and get a first TTS call returning audio.
Switching to or from Voicemaker
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From ElevenLabs: export your script library, re-record each piece through Pulse 1.0 or ProPlus, and rebuild voice profiles using Voicemaker's cloning slots.
- →From PlayHT: move saved scripts into Voicemaker Projects, then re-create any cloned voices with the Pro-plan cloning slots.
- →From a cloud TTS API (Google/Amazon): point your integration at Voicemaker's REST endpoints — the docs cover TTS, STT and STS with MP3, WAV and telephony output formats.
- →From manual narration or freelancers: upload existing PDF/DOC/TXT scripts and generate a first take, then use VoxStudio's Voice Enhancer to close the quality gap.
- ↗To ElevenLabs: export your scripts and credit history, then rebuild each voice with ElevenLabs' voice library and cloning if premium realism becomes the priority.
- ↗To PlayHT: re-upload scripts and re-create any cloned voices, since voice IDs and profiles do not transfer between platforms.
- ↗To a cloud TTS API: replace Voicemaker endpoints with your provider's, and re-map output formats (Voicemaker's WAV 48kHz and 8kHz telephony map cleanly).
- ↗To an open-source engine: if you need offline or on-premise TTS, export your scripts and rebuild voices locally — Voicemaker has no self-hosted option.
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Voicemaker”, and we withheld 5: 5 could not be judged, because “Voicemaker” is a single word that other videos use for other things. Showing the 1 we can prove is about Voicemaker.
Official links
Tools that pair well with Voicemaker
Common stack mates teams adopt alongside Voicemaker, with the specific reason each pairing earns its keep.
AIAI.com
AIAI.com bundles text-to-image, video, voice-cloning, music, and study tools behind one credit-based subscription.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Speechify Studio - AI Voice Generator
AI voice generator with 1,000+ lifelike voices in 60+ languages, plus dubbing, cloning, and avatars
Featured Head-to-Head Comparisons
Voicemaker vs Splice
Splice and Voicemaker serve entirely different needs. Splice is for music producers who need royalty-free samples and rent-to-own plugins; Voicemaker is for content creators and developers needing expressive text-to-speech with multi-language support. Buy based on whether you need samples or voiceovers—they don't compete directly.
Voicemaker vs Landr Mastering
Lanndr and Voicemaker serve completely different needs. Choose LANDR if you're a musician seeking quick, affordable AI mastering with stem control. Choose Voicemaker if you need expressive, multilingual text-to-speech with voice cloning and multi-speaker capabilities. There's no overlap—pick based on your creative output: music or voice.
Voicemaker vs Storyfile
Choose StoryFile if you need authentic, recorded video conversations for historical or legacy exhibits; it’s unmatched in emotional authenticity but enterprise-priced. Choose Voicemaker if you need scalable, multilingual TTS with voice cloning and expressive control—it’s far more flexible and affordable for content creation and development.
Alternatives to Voicemaker
View allAIAI.com
AIAI.com bundles text-to-image, video, voice-cloning, music, and study tools behind one credit-based subscription.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Speechify Studio - AI Voice Generator
AI voice generator with 1,000+ lifelike voices in 60+ languages, plus dubbing, cloning, and avatars
Frequently Asked Questions
Categories
Best-of guides
Used Voicemaker? Help shape our editorial sentiment research.
