Voicebox

Voicebox

Voicebox is an open-source desktop AI voice cloning studio for local cloning, dictation, and agent speech — no account, no cloud, MIT

74/100Safe BetFree · from $12/yearFreemium

If you already have the GPU and care where your voice data lives, Voicebox is the value pick — the free tier is the whole app, not a time-limited trial. The MCP integration is the part cloud TTS vendors can't easily match: one voicebox.speak call hands Claude Code, Cursor, or Cline a cloned voice running locally. The catch is setup friction and the absence of a web or mobile client, so the $12/year Cloud and $48/year Studio plans are placeholders until they actually launch.

Verified 19m ago · liveness 74/100 · cite: rightaichoice.com/tools/voicebox

Best for
  • Content creators who want local voice cloning and multi-voice story editing without a subscription
  • Privacy-conscious users who won't send voice data to a cloud TTS vendor
  • Developers giving MCP-aware agents like Claude Code, Cursor, or Cline a cloned voice
  • Podcasters and storytellers building multi-character audio with a timeline editor
Not ideal for
  • Anyone who needs a web or mobile client today — desktop only for now
  • Teams needing cross-device cloud sync immediately — Cloud and Studio are marked coming soon
  • Beginners without a capable GPU who don't want to wrestle with driver and inference setup
Visit Website

IntermediateAllowing for a GPU driver and inference runtime install (Metal, CUDA, ROCm, Intel Arc, or DirectML), budget roughly 30–60 minutes to first cloned voice; a remote-server path with one-click setup and automatic discovery shortens the GPU wrangling for developers. Dictation is faster to first value — install, pick a Whisper size, and you're speaking into any app within about 15 minutes.DesktopNo public APIVerified 19m ago
Pricing
Free · from $12/year
FreemiumFree tier3 plans3 hidden costs
Learning curve
Intermediate
Allowing for a GPU driver and inference runtime install (Metal, CUDA, ROCm, Intel Arc, or DirectML), budget roughly 30–60 minutes to first cloned voice; a remote-server path with one-click setup and automatic discovery shortens the GPU wrangling for developers. Dictation is faster to first value — install, pick a Whisper size, and you're speaking into any app within about 15 minutes.
Runs on
Desktop
No public API · 3 integrations
Who it's for
Content creatorDeveloper using Claude CodePrivacy-conscious knowledge worker
Live sentiment
Is Voicebox actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Voicebox if you need a web or mobile client, aren't willing to run a capable local GPU, or need an enterprise SLA today — the paid cloud tiers are still marked coming soon.

The 30-second take
Biggest gripe

Cloud and Studio plans are 'Coming soon' and the vendor states pricing and limits are not final, so the $12/year and $48/year figures may change at launch.

Price reality

The full app is $0 forever — that beats ElevenLabs and WisprFlow subscriptions outright for anyone with a capable GPU. The paid tiers are add-on cloud storage: $12/year Cloud (25 GB, 5 devices) and $48/year Studio (250 GB, unlimited devices), both marked coming soon with pricing not final. Solo creators and privacy-first users stay on free; only those wanting encrypted cross-device backup need to pay.

In short

Voicebox — Voicebox is an open-source desktop AI voice cloning studio for local cloning, dictation, and agent speech — no account, no cloud, MIT. Best for Content creators who want local voice cloning and multi-voice story editing without a subscription, Privacy-conscious users who won't send voice data to a cloud TTS vendor, Developers giving MCP-aware agents like Claude Code, Cursor, or Cline a cloned voice. Free to start; paid plans from $12/yr.

What's new in Voicebox

Checked 9 days ago

Across the latest 1 update: 1 news mention.

What people actually say about Voicebox — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

13 mentions across 2 sources (Hacker News, Lemmy) · researched Jul 3, 2026.

43% positive57% critical

Average across the 2 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Free and open-source under MIT license.
  • +Fully local, no account or cloud required.
  • +Voice cloning from just a few seconds of audio.
  • +Supports multiple TTS engines for varied quality.
  • +Unlimited generation up to 50,000 characters per go.
Recurring frustrations
  • −Community data extremely sparse; no real user feedback.
  • −Hands-on reviews are missing—reliability unclear.
  • −No official support; relies on open-source community.
  • −Setup may be complex for non-technical users.
  • −Cloud backup/sync feature is still in development.
Patterns worth knowing
Privacy-first alternative to cloud voice services
Seen on Hacker News
Scarce hands-on user feedback
Seen on Hacker News, Lemmy
Mentioned in context of voice data breaches
Seen on Hacker News, Lemmy
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • Potential hardware cost for GPU acceleration
  • • No paid tier currently; future cloud backup may introduce costs

Viability Score

74/100
Safe Bet

How well maintained and how widely used is Voicebox? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
43
What the vendor publishes
40

Last calculated: October 2026

How we score →

Key Features

  • Voice cloning from as little as 3 seconds of audio
  • Clone from uploaded file (WAV/MP3/FLAC/WebM), microphone (30s max), or system audio capture
  • Seven TTS engines including Qwen and Chatterbox
  • Timeline-based Stories Editor for arranging multi-voice narratives
  • Audio effects pipeline: pitch shift, reverb, delay, compression, savable presets
  • Preview effects live and set defaults per voice profile
  • Dictation anywhere: hold a hotkey, speak, paste into any app
  • Whisper transcription in five sizes (Base, Small, Medium, Large, Turbo), 99 languages
  • Optional local LLM refines transcripts by removing ums and self-corrections
  • MCP integration: voicebox.speak gives any MCP-aware agent a cloned voice
  • Per-agent voice binding so each MCP client speaks in its own profile
  • POST /speak endpoint for ACP, A2A, shell scripts, and custom harnesses
  • Personalities: Rewrite your text or Compose fresh lines in character
  • Generation up to 50,000 characters with sentence-boundary splitting and crossfade
  • Local GPU inference via Metal, CUDA, ROCm, Intel Arc, or DirectML, plus remote server option

About Voicebox

FreemiumIntermediateNo APIDesktop

Voicebox is a free, open-source desktop app for AI voice cloning, text-to-speech, and dictation that runs entirely on your own machine. It's built for people who want professional voice tools without sending audio to a cloud vendor: content creators, podcasters, privacy-focused users, and developers wiring speech into AI agents. Download it for macOS, Windows, or Linux, and you can clone a voice from as little as three seconds of audio pulled from a file, a microphone recording (30 seconds max), or audio captured while it plays on your system — a YouTube video or podcast, say. Generation spans seven TTS engines, including Qwen and Chatterbox, and there's no shortage of things to do with the output. A timeline-based Stories Editor lets you arrange tracks, trim clips, and mix conversations between characters, while an effects pipeline handles pitch shift, reverb, delay, and compression with savable presets per voice profile. Generation runs up to 50,000 characters in one go, auto-split at sentence boundaries and crossfaded. Dictation works anywhere: hold a shortcut, speak, release, and a cleaned transcript lands in the focused text field or your clipboard. Whisper handles speech-to-text across five sizes (Base, Small, Medium, Large, Turbo) in 99 languages, with an optional local LLM that strips ums and self-corrections without rephrasing your words. The MCP layer may be the real draw. One voicebox.speak call gives any MCP-aware agent — Claude Code, Cursor, Cline — a cloned voice, with per-agent profile binding so you know which agent is talking. There's also a POST /speak endpoint for ACP, A2A, and custom harnesses. The app itself is free forever under the MIT license; the project introduced a $VOICEBOX token after passing 1M downloads, with the vendor citing that donations alone are unsustainable. Optional end-to-end encrypted cloud backup and sync sit at $12/year (Cloud) and $48/year (Studio), both still marked coming soon. Versus ElevenLabs or WisprFlow, the

Behind the Verdict

The reason to pick Voicebox is simple: it's the whole app for $0. Cloning, every TTS engine, dictation, the Stories Editor, MCP, personalities — all local, no account. If your machine has a decent GPU, that's a genuinely unusual deal, and it's the reason we'd reach for it over ElevenLabs when we don't need cloud convenience. Where it earns its keep is agent speech. Binding Claude Code to one voice profile and Cursor to another, then hearing each answer out loud through a visible pill, is a small workflow change that makes multi-agent sessions far easier to follow. The POST /speak endpoint covers anything that doesn't speak MCP. Dictation is the sleeper feature for anyone who already writes and chats on a desktop all day. Hold the shortcut, speak, release — the transcript lands wherever your cursor is. Whisper covers 99 languages across five model sizes, and the optional local LLM polish keeps the 'ums' out without rewriting your sentence, which is the failure mode most cleanup tools fall into. Setup is the honest cost. Local GPU inference via Metal, CUDA, ROCm, Intel Arc, or DirectML means drivers, model downloads, and the occasional version mismatch. There's a remote-server option with automatic discovery if you'd rather point the app at a beefier machine, which is the sane answer for laptop users without a capable card. The $VOICEBOX token is a real consideration, not a footnote. It's how the project funds itself after 1M downloads without marketing, and the vendor says donations alone don't cover it. That's a legitimate model, but if you're deploying Voicebox inside a company, check what the token implies for your procurement process before you standardize on it. We'd pass on Voicebox if you need a web or mobile client today, if you want cross-device sync right

Researching Voicebox? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Voicebox actually fits — and what changes day-one when you adopt it.

Content creator

You drop a 10-second clip of your own voice, clone it in Voicebox, then build a multi-character story in the Stories Editor — arranging tracks, trimming clips, and mixing a conversation between your clone and a second profile.

Outcome: A finished multi-voice audio piece, produced locally with no subscription and no audio leaving your machine.

Developer using Claude Code

You wire Voicebox's MCP integration into Claude Code, bind a cloned voice profile per agent, and every time the agent completes a task the voicebox.speak call reads the result aloud through the pill overlay.

Outcome: Agents talk back in voices you own, so you can hear which agent is speaking without reading the terminal.

Privacy-conscious knowledge worker

You hold the dictation shortcut, speak a paragraph, release, and the transcript lands in the focused field — with a local LLM stripping ums and self-corrections before it pastes.

Outcome: Hands-free text entry across any app, with audio kept alongside the transcript and nothing sent to a cloud STT vendor.

Use Cases

Models Under the Hood

Qwen 1.7BQwen 0.6BQwen3ChatterboxWhisper Base (74M)Whisper Small (244M)Whisper Medium (769M)Whisper Large (1.5B)Whisper Turbo (809M)

as of 2026-10-04

Limitations

  • Voicebox is a local-first desktop app: the evidence states it runs entirely on your machine, with no account required, and the paid Cloud and Studio tiers for encrypted backup and sync are still marked coming soon with pricing and limits not final.
  • The free Local plan provides the full app including cloning, dictation, MCP/agent integration, and unlimited local generations at no cost.
  • The site does not publish minimum hardware requirements, so specifics about GPU capability are not confirmed by the provided pages.

as of 2026-09-15

Verification history

We have re-verified Voicebox 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 8 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Voicebox tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Local

$0

Ideal for

Solo creators, privacy-first users, and developers who already have a capable GPU and just want the full local app with no account.

What this tier adds

Free entry point — the entire app: cloning, dictation, every TTS engine, MCP, and personalities, all local.

Cloud

$12/year

Ideal for

Individuals who want encrypted off-device backup and sync across a handful of desktops and mobile devices.

What this tier adds

Adds end-to-end encrypted backup and sync, 25 GB storage, up to 5 devices, and 30-day version history. Coming soon.

Studio

$48/year

Ideal for

Power users and small professionals with larger audio libraries who need unlimited devices and priority support.

What this tier adds

Steps storage to 250 GB, lifts the device cap, extends version history to 1 year, and adds priority support. Coming soon; placeholder pricing TBD.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Cloud and Studio plans are 'Coming soon' and the vendor states pricing and limits are not final, so the $12/year and $48/year figures may change at launch.
  • Cloud backup starts at $12/year for 25 GB and 5 devices; heavier users needing 250 GB and unlimited devices step up to Studio at $48/year.
  • You supply the GPU and the electricity — local Whisper Large and TTS inference are hardware-intensive, and a capable GPU is a real upfront cost.

Where the pricing makes sense

The company stage and team size where Voicebox's pricing actually pencils out — and where peers do it cheaper.

The full app is $0 forever — that beats ElevenLabs and WisprFlow subscriptions outright for anyone with a capable GPU. The paid tiers are add-on cloud storage: $12/year Cloud (25 GB, 5 devices) and $48/year Studio (250 GB, unlimited devices), both marked coming soon with pricing not final. Solo creators and privacy-first users stay on free; only those wanting encrypted cross-device backup need to pay.

Setup time & first value

How long it actually takes to get something useful out of Voicebox — broken out by persona, not the marketing-page minute.

Allowing for a GPU driver and inference runtime install (Metal, CUDA, ROCm, Intel Arc, or DirectML), budget roughly 30–60 minutes to first cloned voice; a remote-server path with one-click setup and automatic discovery shortens the GPU wrangling for developers. Dictation is faster to first value — install, pick a Whisper size, and you're speaking into any app within about 15 minutes.

Switching to or from Voicebox

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From ElevenLabs: Export your reference clips, then create Voicebox profiles from the same audio and re-generate locally.
  • →From WisprFlow: Move dictation to Voicebox's hotkey flow, choosing Whisper Base through Turbo to match your hardware.
  • →From cloud TTS tools generally: Keep your source voice samples, import them as Voicebox profiles, and reproduce generation locally.
Migrating out
  • ↗To ElevenLabs: Export your Voicebox-generated audio from the local library and upload reference clips to ElevenLabs for cloud generation.
  • ↗To a cloud STT service: Run your Voicebox captures through a hosted transcription API if you later need server-side processing.

Integrations

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Voicebox”, and we withheld 6: 6 could not be judged, because “Voicebox” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Voicebox.

Tools that pair well with Voicebox

Common stack mates teams adopt alongside Voicebox, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Voicebox

View all
OmniVoice Studio

OmniVoice Studio

Open-source desktop studio for voice cloning, voice design, dubbing, dictation and audiobooks — runs on your own machine.

FreemiumTry
Fish Audio

Fish Audio

Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.

FreemiumTry
ElevenLabs

ElevenLabs

ElevenLabs turns text into ultra-realistic AI voice, music, dubbing and conversational voice agents from one credit pool.

FreemiumTry

Frequently Asked Questions

Used Voicebox? Help shape our editorial sentiment research.