Hume AI
Hume AI provides real human feedback, simulation, and expression measurement for emotionally intelligent voice AI.
Pick Hume AI when the thing you need to measure is emotional quality and naturalness — the dimensions automated metrics miss. Its Human Feedback API returns per-sample human scores in hours, Kairos generates and replays scenarios at scale, and the Expression Measurement API gives you 48+ emotion categories and 600+ voice descriptors as a real-time signal. The generation side is genuinely useful too: Octave 2 (preview) and EVI 3/EVI 4 mini, with turn detection and interruption controls and an external LLM path covering claude-opus-4-6, gpt-5.1, and gpt-5.2 variants. Choose ElevenLabs instead if raw TTS fidelity is the whole job and you don't need human evaluation. Choose Vapi or Retell if
Verified 1d ago · liveness 87/100 · cite: rightaichoice.com/tools/hume-ai
- Voice AI teams that need human-grounded evaluation, not just accuracy metrics
- Developers building emotionally aware voice assistants and speech-to-speech agents
- Researchers needing annotated speech datasets and human rating studies
- Teams regression-testing voice models against real-world scenarios
- Teams whose only requirement is raw TTS fidelity with no evaluation layer
- Developers who need a fully open-source voice stack
- Projects with no emotional-nuance requirement — basic TTS or STT is simpler elsewhere
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Hume AI if you only need a voice to read text and have no interest in measuring how that voice lands with real listeners.
Going past your tier's character allowance is billed per 1,000 characters — $0.15 on Creator down to $0.05 on Business — so a viral month costs more than the listed monthly price.
The free $0 tier and the $3/mo Starter exist mainly to test the API — 30,000 TTS characters and 40 EVI minutes go fast. Solo developers generally land on $7/mo Creator or $70/mo Pro, while $200/mo Scale and $500/mo Business fit teams that need 150-225 RPM and bundled seats. ElevenLabs competes on raw TTS fidelity at similar price points; Hume's differentiator is the human evaluation layer you don't get there.
In short
Hume AI — Hume AI provides real human feedback, simulation, and expression measurement for emotionally intelligent voice AI. Best for Voice AI teams that need human-grounded evaluation, not just accuracy metrics, Developers building emotionally aware voice assistants and speech-to-speech agents, Researchers needing annotated speech datasets and human rating studies. Free to start; paid plans from $3/mo.
What's new in Hume AI
Checked yesterdayAcross the latest 2 updates: 2 news mentions.
Google's Gemini 3.8 Flash TTS tops Hume's Real-World VoiceEQ leaderboard
Hume reports that Google's Gemini 3.8 Flash TTS took the top spot on its Real-World VoiceEQ leaderboard, a multidimensional benchmark spanning TTS, speech-to-speech, speech understanding, and ASR.
Emotional Intelligence Is a Training-Time Property, Not a Prompt
Hume argues that emotional intelligence in voice models is set during training rather than by prompting, framing EQ as a model-level property rather than something you can instruct your way into.
Viability Score
How well maintained and how widely used is Hume AI? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Real-time Expression Measurement API covering 48+ emotion categories and 50+ languages
- 600+ output metrics and voice descriptors for expression analysis
- Kairos simulation platform for agent-to-agent and human-to-agent conversation testing
- Human Feedback API with per-sample scores, free-response feedback, and aggregated analysis
- Pre-screened human raters with fraud detection and quality monitoring
- Text-to-speech with Octave 1 and Octave 2 (preview) speech-language models
- Speech-to-speech with EVI 3 and EVI 4 mini models
- Configurable turn detection and interruption settings per EVI config (April 2026)
- End-of-turn silence tuning from 500ms to 3000ms
- Experimental temperature parameter for TTS output variation (May 2026)
- External LLM support: claude-opus-4-6, gpt-5.1, gpt-5.1-priority, gpt-5.2, gpt-5.2-priority
- ZERO prompt expansion mode for full system prompt control
- Voice cloning - create and use unlimited on all tiers, API access on Enterprise
- Voice library of over 100 expressive voices plus voice design from prompts
- SDKs for React, TypeScript, Python, Swift, and .NET
About Hume AI
Hume AI is the data and evaluation layer for emotionally intelligent voice AI — it helps you build and measure voice models the way people actually experience them. The platform has four parts. Data Solutions collects custom speech data built around your own scenarios, use cases, and evaluation requirements. Kairos is a simulation platform where you build evaluation suites from real-world use cases, run agent-to-agent and human-to-agent conversations, and track regressions over time. The Expression Measurement API returns real-time and offline analysis across 48+ emotion categories and 50+ languages with 600+ voice descriptors. The Human Feedback API routes work to pre-screened human raters who return per-sample scores, free-response feedback, and aggregated analysis — in hours rather than days. On the generation side, Hume ships Octave 2 (preview) and Octave 1 for expressive text-to-speech built on a speech-language model rather than conventional read-the-words TTS, plus EVI 3 and EVI 4 mini for real-time speech-to-speech, with an external LLM path supporting claude-opus-4-6, gpt-5.1, gpt-5.1-priority, gpt-5.2, and gpt-5.2-priority. Hume also publishes two leaderboards: Real World VoiceEQ Bench, which ranks voice models across recognition, understanding, expression, and conversation by human judgment, and SLM Judge, which scores automated evaluators against human ratings. It fits voice AI teams, researchers, and developers who need human-grounded measurement of emotional quality, not just accuracy metrics.
Behind the Verdict
Hume AI is unusual in the voice market because it sells the measurement problem first and the speech models second. The four-product structure is coherent: Data Solutions for collection, Kairos for simulation, Expression Measurement for real-time signal, Human Feedback for ground truth. For a team running model development, that's the shape of a real evaluation loop — generate scenarios, run them, score against humans, catch regressions, repeat. The claim that ratings come back "in hours, not days" is the core value proposition, because human eval is normally the bottleneck that forces teams onto automated judges of uncertain trustworthiness. Hume's SLM Judge leaderboard is a candid answer to that problem: instead of arguing that LLM judges are good enough, it publishes which judges track human ratings most closely, so you can pick one you can defend. On generation, Octave is positioned as a speech-language model rather than a TTS reader, and the roadmap has moved in the direction of operator control. February 2026 added supplemental LLM support (claude-opus-4-6, gpt-5.1, gpt-5.1-priority, gpt-5.2, gpt-5.2-priority) plus a ZERO prompt expansion mode for full system prompt control. April 2026 exposed per-config turn detection and interruption settings — end-of-turn silence between 500 and 3000ms, speech detection threshold, prefix padding, and minimum interruption duration. May 2026 added an experimental temperature parameter to TTS endpoints, higher for variation, lower for consistency. These are the knobs that separate a demo from a production voice agent, and they arrived in the last six months. Two honest constraints. First, pricing is usage-based underneath the seat price — every tier has an additional-characters rate and an additional EVI 3 rate, so a fixed monthly number is a starting point, not a ceiling. Second, Hume's own research position, argued in July 2026, is that emotional intelligence is a training-time property rather than a prompt-time one — which is a strong claim, but it also means you should not expect prompt engineering to close an expressiveness gap the model doesn't have. If raw fidelity is your only axis, compare ElevenLabs directly. If you need to know whether callers are expressing frustration and whether your model is getting better at handling it, Hume is built for exactly that question.
Researching Hume AI? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Hume AI actually fits — and what changes day-one when you adopt it.
You pull last quarter's real support calls, use Kairos to turn the recurring failure moments into an automated evaluation suite, then run the updated agent against it after each model change.
Outcome: Regressions surface before release rather than in a customer escalation, and you have a scenario library that keeps growing.
You run your candidate models through the Expression Measurement API for real-time vocal signal, then route the same clips to the Human Feedback API for per-sample human scores.
Outcome: You get a side-by-side ranking grounded in human judgment plus the automated metrics that explain where each model diverges.
You start on the free tier, wire up EVI 4 mini with the TypeScript SDK, tune turn detection and interruption settings, and add Octave narration.
Outcome: A working emotionally responsive prototype before you spend anything, then a move to Creator once real usage starts.
Use Cases
- Measure whether a call-center voice agent reduces caller frustration using the Expression Measurement API
- Run Human Feedback API studies to score TTS and speech-to-speech models against human judgment
- Build Kairos evaluation suites from real support transcripts and replay them as regression tests
- Pick a trustworthy automated evaluator using the SLM Judge leaderboard instead of guessing
- Generate expressive audiobook and podcast narration with Octave 1 and Octave 2
- Prototype an emotionally aware coaching or interview simulator with EVI 3 turn detection controls
- Run agent-to-agent conversations at scale to find failure modes before real callers hit them
- Add real-time vocal expression signals to a speech-to-speech assistant built on gpt-5.2 or claude-opus-4-6
Models Under the Hood
as of 2026-09-22
Limitations
- Every paid tier is usage-based underneath the seat price: additional TTS characters run $0.15, $0.12, $0.10, or $0.05 per 1,000 depending on tier, and additional EVI 3 minutes run $0.06, $0.05, or $0.04 per minute.
- Included allowances are small at the bottom — Free gives 10,000 TTS characters (~10 minutes) and 5 EVI minutes at 15 RPM with 1 concurrent external-LLM connection, and Creator gives 140,000 characters and 200 EVI minutes.
- Requests per minute climb in steps (15, 15, 75, 75, 150, 225) so high-throughput workloads need a higher tier.
- Team seats are bundled rather than free — 3 on Scale, 5 on Business, unlimited only on Enterprise.
- API access for voice cloning, custom RPM and overage pricing, and SOC 2 Type II, GDPR, and HIPAA compliance are Enterprise-only; support is Discord-only below Business.
as of 2026-09-28
Verification history
We have re-verified Hume AI 18 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 18 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Hume AI tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0/mo
Ideal for
Developers kicking the tires on the API and the no-code playground before committing budget
What this tier adds
Free entry point: 10,000 TTS characters (~10 minutes), 5 EVI minutes, 15 RPM, 1 concurrent external-LLM connection, Discord support
Starter
$3/mo
Ideal for
Solo builders past the free allowance who still need only a few minutes of voice per month
What this tier adds
Adds 20,000 more TTS characters and 35 more EVI minutes over Free, plus 5 concurrent external-LLM connections
Creator
$7/mo (first month 50% off, $14/mo after)
Ideal for
Indie developers shipping a narration or voice feature that real users touch
What this tier adds
Jumps to 140,000 TTS characters, 200 EVI minutes, 75 RPM, and a $0.15/1,000 overage rate
Pro
$70/mo
Ideal for
Small teams running a production voice agent who need Serious monthly volume
What this tier adds
Adds 1,000,000 TTS characters, 1,200 EVI minutes, 10 concurrent connections, and a lower $0.12/1,000 overage
Scale
$200/mo
Ideal for
Growing voice teams that need throughput, collaboration, and cheaper overage rates
What this tier adds
Raises to 3,300,000 characters, 5,000 EVI minutes, 150 RPM, 20 concurrent connections, and includes 3 team seats
Business
$500/mo
Ideal for
Established companies running high-volume voice workloads with a support SLA expectation
What this tier adds
Adds 10,000,000 characters, 12,500 EVI minutes, 225 RPM, 30 concurrent connections, 5 team seats, and Slack support
Enterprise
Custom
Ideal for
Regulated or very-high-volume organizations that need compliance documentation and negotiated terms
What this tier adds
Adds unlimited usage, custom RPM and overage pricing, voice cloning API access, and SOC 2 Type II, GDPR, and HIPAA compliance
Where the pricing makes sense
The company stage and team size where Hume AI's pricing actually pencils out — and where peers do it cheaper.
The free $0 tier and the $3/mo Starter exist mainly to test the API — 30,000 TTS characters and 40 EVI minutes go fast. Solo developers generally land on $7/mo Creator or $70/mo Pro, while $200/mo Scale and $500/mo Business fit teams that need 150-225 RPM and bundled seats. ElevenLabs competes on raw TTS fidelity at similar price points; Hume's differentiator is the human evaluation layer you don't get there.
Setup time & first value
How long it actually takes to get something useful out of Hume AI — broken out by persona, not the marketing-page minute.
With an API key and the TypeScript or Python SDK, a first TTS or EVI call is roughly an afternoon — the docs include quickstarts and example repos. Wiring turn detection, interruption settings, and external LLM configuration into a production agent is a few days. A full Kairos evaluation suite with human raters on top typically runs one to two weeks, mostly because scenario design and rater
Switching to or from Hume AI
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From ElevenLabs: keep ElevenLabs for pure narration if it tests better and route evaluation plus emotional voice work to Hume.
- →From raw OpenAI or Anthropic speech APIs: point the external LLM config at claude-opus-4-6 or gpt-5.2 and add Hume's expression measurement on top.
- →From a homegrown human-eval spreadsheet: move the rubric into Kairos scenarios and the scoring into the Human Feedback API.
- →From Vapi or Retell: keep the orchestration layer and call Hume's APIs for measurement rather than swapping platforms outright.
- ↗To ElevenLabs: if your only remaining need is raw TTS fidelity and the evaluation layer is no longer earning its cost.
- ↗To Vapi or Retell: if you decide you want a turnkey agent platform instead of assembling voice plus measurement yourself.
- ↗To an in-house evaluation stack: only if you are prepared to recruit, screen, and QA your own human rater pool.
Integrations
Resources & Guides
- Documentationhume.ai
Welcome to Hume AI | Hume API
Hume AI builds AI models that enable technology to communicate with empathy and support human well-being.
- Documentationhume.ai
Welcome to Hume AI | Hume API
Hume AI builds AI models that enable technology to communicate with empathy and support human well-being.
- Quickstarthume.ai
Welcome to Hume AI | Hume API
Hume AI builds AI models that enable technology to communicate with empathy and support human well-being.
- Documentationhume.ai
Welcome to Hume AI | Hume API
Hume AI builds AI models that enable technology to communicate with empathy and support human well-being.
- Documentationhume.ai
Support | Hume API
Get technical support, contact our team, or explore enterprise and research programs.
- Documentationhume.ai
Welcome to Hume AI | Hume API
Hume AI builds AI models that enable technology to communicate with empathy and support human well-being.
- Guidehume.ai
Welcome to Hume AI | Hume API
Hume AI builds AI models that enable technology to communicate with empathy and support human well-being.
- Resourcehume.ai
The AI toolkit for voice and emotion
Building AI with emotional intelligence to create technology that truly understands humanity.
Tutorials & Learning
YouTube returned 6 videos for “Hume AI”, and we withheld 6: 6 could not be judged, because “Hume AI” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Hume AI.
Official links
Tools that pair well with Hume AI
Common stack mates teams adopt alongside Hume AI, with the specific reason each pairing earns its keep.
Hume AI Octave 2
Emotionally expressive text-to-speech and speech-to-speech voice AI with human-judged evaluation built in.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Speechify Studio - AI Voice Generator
AI voice generator with 1,000+ lifelike voices in 60+ languages, plus dubbing, cloning, avatars, and captions
Alternatives to Hume AI
View allHume AI Octave 2
Emotionally expressive text-to-speech and speech-to-speech voice AI with human-judged evaluation built in.
Fish Audio
Fish Audio turns text into expressive, emotionally controllable speech with voice cloning from 15 seconds of audio and a free developer TTS API.
Speechify Studio - AI Voice Generator
AI voice generator with 1,000+ lifelike voices in 60+ languages, plus dubbing, cloning, avatars, and captions
Frequently Asked Questions
Best-of guides
Used Hume AI? Help shape our editorial sentiment research.