ComfyUI VoxCPM
Open-source multilingual text-to-speech with voice cloning and voice design, running locally inside your own ComfyUI workflow.
If your pipeline already lives in ComfyUI, VoxCPM2 removes the hop out to a hosted speech API — cloning, voice design and LoRA tuning all happen locally on Apache-2.0 weights you keep. The 30-language tag set is the widest thing here that competitors rarely match at this price, and the tokenizer-free design (arXiv:2509.24650) is a genuine architectural choice, not a rebrand. It is not a latency play and not a product you install and forget: budget a GPU that holds a ~4.6GB BF16 checkpoint and real setup time. Pick it for control and privacy; pass on it in favour of ElevenLabs or Azure Speech if you need sub-second streaming or a managed endpoint.
Verified 6d ago · liveness 75/100 · cite: rightaichoice.com/tools/comfyui-voxcpm
- ComfyUI users who want TTS as just another node in an existing generation graph
- Game and narrative teams building character voices with emotion and prosody control
- Creators producing multilingual voiceovers without sending scripts to a cloud service
- Researchers working on diffusion TTS and tokenizer-free speech architectures
- Real-time voice agents and phone bots that need low-latency streaming
- Anyone wanting a plug-and-play API key and zero environment setup
- Teams without a GPU that can comfortably hold a ~4.6GB BF16 model
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip VoxCPM2 if you need sub-second streaming for a live voice agent, or you want a hosted endpoint and an API key instead of provisioning a GPU that can hold a ~4.6GB BF16 checkpoint.
You pay in hardware, not invoices: the ~4.6GB BF16 checkpoint needs a GPU that holds it comfortably, and that capacity is yours to fund.
VoxCPM2 is free to download under Apache-2.0 — there is no per-character or per-seat line item here. Solo creators with a capable GPU spend nothing but time; teams without one should compare the cost of the GPU instance they'd rent against hosted speech services like ElevenLabs or Azure Speech, which bill per character or per hour and include the endpoint, uptime and support you'd otherwise build yourself.
In short
ComfyUI VoxCPM — Open-source multilingual text-to-speech with voice cloning and voice design, running locally inside your own ComfyUI workflow. Best for ComfyUI users who want TTS as just another node in an existing generation graph, Game and narrative teams building character voices with emotion and prosody control, Creators producing multilingual voiceovers without sending scripts to a cloud service. Free to use.
What's new in ComfyUI VoxCPM
Checked 6 days agoAcross the latest 2 updates: 1 launch and 1 changelog entry.
openbmb/VoxCPM2 last modified
The VoxCPM2 model repository was last updated on 18 August 2026, with 15 open discussions listed on the card.
openbmb/VoxCPM2 model card created
The VoxCPM2 repository was created on Hugging Face, publishing Apache-2.0 safetensors weights at 2,290,004,544 BF16 parameters for text-to-speech, voice cloning and voice design.
What people actually say about ComfyUI VoxCPM — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
26 mentions across 2 sources (YouTube, GitHub) · researched Jul 5, 2026.
Average across the 2 sources that answered — each source counts once, not each post.
- +Free alternative to premium TTS like ElevenLabs.
- +Supports 30 languages with multilingual TTS.
- +High-quality 48kHz audio output praised by users.
- +Controllable emotion and prosody in voice cloning.
- +Diffusion-based model produces natural-sounding speech.
- −Installation errors and missing nodes frustrate beginners.
- −Continuation cloning + voice design produces corrupted output.
- −Version 1.6 has audio pitch errors and word drops.
- −Limited node availability in ComfyUI integration.
- −No native LoRA loading for custom voices yet.
- • Requires GPU hardware with sufficient VRAM
- • Time spent troubleshooting installation and bugs
Viability Score
How well maintained and how widely used is ComfyUI VoxCPM? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- 30-language multilingual text-to-speech
- Voice cloning from a short reference clip with controllable similarity
- Voice design from a text prompt with no reference audio
- LoRA training to fine-tune a voice toward a character or style
- Tokenizer-free architecture for context-aware speech generation
- Diffusion-based speech generation
- Emotion and prosody control on generated speech
- Runs fully locally with audio kept on-device
- Node-based TTS workflow inside ComfyUI
- Safetensors weights at 2,290,004,544 BF16 parameters (~4.6GB, unsharded)
- Apache-2.0 license permitting commercial use and modification
- Tagged for SageMaker deployment
- Language coverage including Arabic, Burmese, Khmer, Lao, Swahili and Tagalog
- Peer-reviewed methodology via arXiv paper 2509.24650
- Active Hugging Face repo with 2.66M+ all-time downloads and 1,638 likes
About ComfyUI VoxCPM
ComfyUI VoxCPM is the ComfyUI-facing node for VoxCPM2, an Apache-2.0 diffusion text-to-speech model published by OpenBMB at openbmb/VoxCPM2 on Hugging Face. You run it on your own GPU and wire speech synthesis into a node-based generation graph rather than calling a hosted API. The repo's tag set lists 30 languages — Chinese, English, Arabic, Hindi, Japanese, Korean, Thai, Vietnamese, Swahili, Burmese, Khmer, Lao, Tagalog and Malay alongside Danish, Dutch, Finnish, French, German, Greek, Hebrew, Indonesian, Italian, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish and Turkish. The published methodology is arXiv:2509.24650, "VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning," which is where the tokenizer-free claim comes from. Two workflows carry most of the load. Voice cloning takes a short reference clip and reproduces that speaker with adjustable similarity; voice design skips reference audio entirely and builds a speaker from a text prompt. From there you can train a LoRA adapter toward a character or delivery style, and steer emotion and prosody instead of accepting one fixed read. The weights are safetensors at 2,290,004,544 BF16 parameters — about 4.6GB in a single unsharded file — and the card carries a deploy:sagemaker tag. Traction on the Hub is real: roughly 2.66M all-time downloads and 1,638 likes as of the scrape, with the repo last modified August 2026 and 15 open discussions. Treat it as infrastructure, not a service. You supply the GPU, the environment and the patience, because diffusion speech generation is not a low-latency streaming path. If you want a hosted endpoint and an API key, ElevenLabs or Azure Speech are the shorter route. VoxCPM2 earns its place when the workflow, the weights and the finished audio all need to stay on your machine.
Behind the Verdict
VoxCPM2 sits in the small category of speech models that are genuinely pleasant to own and genuinely annoying to operate. The pleasant part is the scope of what lands in a single Apache-2.0 checkpoint: 30 languages in the tag set, voice cloning with adjustable similarity from a short reference clip, voice design that builds a speaker from a text prompt with no reference audio at all, LoRA training to push a voice toward a character or delivery style, and emotion and prosody control on the output. Nothing about the audio has to leave your hardware, which matters if you are producing under NDA or against data-residency rules. The annoying part is everything around it. VoxCPM2 is distributed as safetensors at 2,290,004,544 BF16 parameters, roughly 4.6GB in one unsharded file, tagged for SageMaker deployment. There is no managed hosting story in what's published — availableInferenceProviders is empty on the model card — so you own the environment, the VRAM and the troubleshooting. The architecture is diffusion-based, and tokenizer-free diffusion speech is a quality-and-control path, not a streaming path. If your requirement is a real-time voice agent, this is the wrong tool and no amount of tuning fixes that; the design goal is fidelity, not milliseconds. Where it fits well: ComfyUI users who want TTS as just another node in a graph they already maintain, game and narrative teams who need many distinct character voices with emotion control, creators doing multilingual voiceovers, and researchers who want to work on tokenizer-free speech architectures with weights they can inspect and modify. The Hub numbers back the adoption: 2.66M+ all-time downloads and 1,638 likes, with the repo last touched August 2026 and 15 open discussions — an active card, not a demo drop. Where it doesn't fit: quick one-off voiceovers, where installing a 4.6GB checkpoint costs more time than it saves; teams without a GPU that can comfortably hold that checkpoint; and anyone who wants to buy inference rather than own it. Compared with ElevenLabs or Azure Speech, you are trading someone else's uptime and API key for total control and zero per-character billing. That trade is worth it for exactly the people described above and wrong for everyone else. One caution on the surrounding ecosystem: the Hugging Face Hub changelog entries in this scrape — live resource usage on Jobs, LeRobot episode previews, Google Cloud Marketplace billing, granular feature access — are platform changes, not VoxCPM model release notes. Don't read them as a VoxCPM roadmap.
Researching ComfyUI VoxCPM? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas ComfyUI VoxCPM actually fits — and what changes day-one when you adopt it.
You load the VoxCPM2 checkpoint, drop the TTS node into the graph that already renders your frames, clone the lead voice from a 20-second reference clip and generate each line with prosody adjustments.
Outcome: Dialogue lands in the same local pipeline as the visuals, and the reference audio and finished voice track never leave your machine.
You use voice design from text prompts to build distinct speakers for each character, then train a LoRA adapter to keep one character's delivery consistent across hundreds of lines.
Outcome: You ship a full cast of voices from one Apache-2.0 model without per-character billing, at the cost of running your own generation pass for every line change.
You render the same script and the same cloned voice across the languages tagged on the model card, including Arabic, Hindi, Swahili and Vietnamese, on a workstation GPU.
Outcome: One voice, many markets, no scripts sent to a cloud service — but you wait on local diffusion generation rather than an API round-trip.
Use Cases
- Create multilingual voiceovers for educational videos across the 30-language tag set.
- Clone a specific speaker's voice for personalized audiobook narration.
- Design unique character voices for indie game dialogue with prosody control.
- Train a LoRA adapter to adapt the model to a custom voice or delivery style.
- Generate studio-quality speech for podcast production on your own GPU.
- Keep script and voice data on-device for work under NDA or data-residency rules.
- Research tokenizer-free diffusion TTS using weights you can inspect and modify.
Models Under the Hood
as of 2026-09-23
Limitations
- VoxCPM2 is an open-source text-to-speech model (Apache-2.0) distributed as Hugging Face safetensors weights at 2,290,004,544 BF16 parameters (~4.6GB, one unsharded file), designed to run locally inside a ComfyUI workflow.
- Diffusion-based generation and LoRA voice fine-tuning require local computational resources — you are running inference, not buying it.
- The model card lists no available inference providers, so plan on provisioning your own GPU environment.
- The repo is actively used, with 2.66M+ all-time downloads and 1,638 likes, and it carries 15 open discussions as of the scrape.
as of 2026-10-02
Verification history
We have re-verified ComfyUI VoxCPM 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 7 verification passes.
Free to cite with attribution — this page re-verifies continuously.
Where the pricing makes sense
The company stage and team size where ComfyUI VoxCPM's pricing actually pencils out — and where peers do it cheaper.
VoxCPM2 is free to download under Apache-2.0 — there is no per-character or per-seat line item here. Solo creators with a capable GPU spend nothing but time; teams without one should compare the cost of the GPU instance they'd rent against hosted speech services like ElevenLabs or Azure Speech, which bill per character or per hour and include the endpoint, uptime and support you'd otherwise build yourself.
Setup time & first value
How long it actually takes to get something useful out of ComfyUI VoxCPM — broken out by persona, not the marketing-page minute.
ComfyUI users already running local models can expect roughly half a day to first audio: download the ~4.6GB unsharded safetensors checkpoint, wire the TTS node into a graph and run a short reference clip. Teams without an existing local inference environment should budget a day or more for driver, CUDA and dependency work before the first line of speech comes out.
Switching to or from ComfyUI VoxCPM
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From a hosted TTS API (ElevenLabs, Azure Speech): export your scripts and produce or source a reference clip for each voice, then clone it locally so you own the model and stop paying per character.
- →From another local TTS model in ComfyUI: swap the TTS node for the VoxCPM2 node, point it at the VoxCPM2 checkpoint and reuse your existing graph wiring.
- →From recording voice talent per project: capture a short reference clip per speaker, clone it once, and regenerate takes on demand instead of booking studio time.
- ↗To a hosted TTS service (ElevenLabs, Azure Speech): take your scripts and voice references to an API that returns audio over HTTP when latency matters more than ownership.
- ↗To a different open TTS checkpoint: keep your ComfyUI graph and reference clips, and swap the model weights if another architecture fits your latency budget better.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “ComfyUI VoxCPM”, and we withheld 6: 6 did not mention ComfyUI VoxCPM. We are showing none, because we could not prove any of them are about ComfyUI VoxCPM.
Official links
Tools that pair well with ComfyUI VoxCPM
Common stack mates teams adopt alongside ComfyUI VoxCPM, with the specific reason each pairing earns its keep.
OmniVoice Studio
Open-source desktop studio for voice cloning, voice design, dubbing, dictation and audiobooks — runs on your own machine.
VoxCPM
Tokenizer-free open-source TTS by OpenBMB: 30 languages, Voice Design, and 48kHz voice cloning you run on your own GPU.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Featured Head-to-Head Comparisons
Comfyui Voxcpm vs Splice
If you need a free, open-source TTS with voice cloning and deep customization in ComfyUI, VoxCPM is unbeatable. For music producers seeking royalty-free samples and rent-to-own plugins with DAW integration, Splice is the clear choice. They serve entirely different creative workflows.
Comfyui Voxcpm vs Storyfile
ComfyUI VoxCPM is the free, open-source choice for developers and creators who need multilingual synthetic audio with voice cloning and LoRA adaptation, especially within ComfyUI workflows. StoryFile is the premium option for museums and legacy projects requiring authentic video-based conversational AI from real people, as demonstrated by recent exhibits with George Takei and Kara Swisher’s CNN digital twin. Your decision hinges on whether you need high-fidelity synthetic audio (VoxCPM) or authentic video interaction (StoryFile).
Comfyui Voxcpm vs Landr Mastering
Choose ComfyUI VoxCPM if you need free, local, multilingual voice synthesis with cloning and deep customization inside ComfyUI. Choose LANDR Mastering for instant, professional AI mastering with a polished plugin and subscription – it’s the best bet for musicians who want release-ready audio without technical overhead. The two tools don’t overlap; pick based on whether you generate voice or master mixes.
Alternatives to ComfyUI VoxCPM
View allOmniVoice Studio
Open-source desktop studio for voice cloning, voice design, dubbing, dictation and audiobooks — runs on your own machine.
VoxCPM
Tokenizer-free open-source TTS by OpenBMB: 30 languages, Voice Design, and 48kHz voice cloning you run on your own GPU.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Frequently Asked Questions
Categories
Best-of guides
Used ComfyUI VoxCPM? Help shape our editorial sentiment research.