VoiceMem
Open-source dual-brain memory for real-time voice agents — facts in the left brain, emotion in the right, streaming at 134ms.
If you're building a voice agent and your memory layer is either flat transcript retrieval or an expensive managed service, VoiceMem is the most interesting open-source option right now — the dual-brain split and streaming prefetch are real architectural choices, not marketing framing. The published LoCoMo (91.2% vs Mem0's 61.68%) and PersonaMem (69.44%) results back it up, and the 134 ms figure is measured against Mem0's 1,440 ms. Weigh it against Mem0 or Zep if you need a managed backend with support. Budget for self-hosting and v0.0.2 rough edges.
Verified 8d ago · liveness 67/100 · cite: rightaichoice.com/tools/voicemem
- Developers building real-time voice agents that need low-latency, emotionally aware persistent memory
- Researchers studying voice AI memory architectures, with an open technical report, eval scripts and ChatMem-400K
- Teams self-hosting an open-source memory backend who want every layer swappable
- Prototypes and MVPs where per-turn token cost is a critical constraint
- Teams needing production SLAs, vendor support or a managed service — this is a v0.0.2 research project
- Non-technical users wanting a hosted API and no self-hosting
- Projects requiring ready-made multi-platform client apps
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip VoiceMem if you need a managed memory backend with an SLA — it's a v0.0.2 open-source library you install, warm up, and debug yourself.
The memory system itself is free under Apache-2.0, but you pay for the LLM API key (openai_key is passed at construction) that drives fact extraction and tagging.
VoiceMem is free under Apache-2.0, which puts it below managed alternatives like Mem0 and Zep (typically per-request or per-seat SaaS pricing) but shifts cost to your infrastructure and engineering hours. It's best when token economics are the binding constraint: at ~430 memory tokens per turn vs Mem0's ~6,956, a high-volume voice product can save real money on LLM calls. If you need vendor support and an SLA, a paid managed service is the right comparison.
In short
VoiceMem — Open-source dual-brain memory for real-time voice agents — facts in the left brain, emotion in the right, streaming at 134ms. Best for Developers building real-time voice agents that need low-latency, emotionally aware persistent memory, Researchers studying voice AI memory architectures, with an open technical report, eval scripts and ChatMem-400K, Teams self-hosting an open-source memory backend who want every layer swappable. Free to use.
What's new in VoiceMem
Checked todayAcross the latest 5 updates: 4 launches and 1 changelog entry.
v0.0.2 — 修复事件日期链路,移除右脑冗余记忆类别,开放可插拔语音合成层
Fixed the event-date chain, removed a redundant right-brain memory category, and opened a pluggable speech-synthesis layer.
v0.0.1 — 发布初代 VoiceMem 和 Technical Report
First VoiceMem release, published alongside the project's technical report.
开源 VoiceMem 模型系列,可直接读取并理解 VoiceMem 提供的记忆
VoiceMem model families open-sourced; they can directly read and understand memories produced by the VoiceMem pipeline.
发布 VoiceMem Utils,开箱即用
VoiceMem Utils package released for out-of-the-box usage.
开源 ChatMem-400K 数据集
ChatMem-400K dataset open-sourced for fine-tuning VoiceMem models on conversational data.
What people actually say about VoiceMem — is it worth it?
We scanned public community sources for VoiceMem on Sep 21, 2026 and could not establish that the discussion we found is about this tool rather than something else sharing its name. Only 2 of the posts we fetched could be positively tied to VoiceMem. Rather than publish a sentiment score built on the wrong subject, we publish nothing here and re-run the scan.
Viability Score
How well maintained and how widely used is VoiceMem? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Dual-brain memory: left brain stores factual schemas and entities, right brain stores persona, emotion, relationships
- Fully streaming pipeline: audio segmentation, ASR, memory extraction, and graph writes while the user speaks
- Speculative prefetching inside a voice turn (0–300 ms) so retrieval starts before the user finishes
- Top-K memory routing and ranking controls context length
- ~300 tokens per single-turn query (project benchmark ~430 memory tokens per turn)
- Published 134 ms response time vs Mem0's 1,440 ms
- 91.2% on LoCoMo with Top-5 memories (Mem0: 61.68%) and 69.44% on PersonaMem
- Multi-modal memory from real audio: voice, speaker, sound events, multi-party conversations, music
- Built-in ASR, speaker verification, scene detection, emotion recognition, and local embedding modules
- Swappable components including the underlying memory engine and a pluggable TTS layer
- SessionBuffer isolates per-session context, with temporary conversations purged at session end
- Two-stage barge-in: VAD pauses and preserves audio queue, clears on stop or stable ASR text
- PCM-sample-based output timeline with AudioWorklet render progress for interrupt handling
- VoiceMem official model families (fine-tuned Qwen reply model) that read WireMem memories
- ChatMem-400K dataset plus finetune pipeline and evaluation scripts for custom training
About VoiceMem
VoiceMem is an open-source memory layer for real-time voice agents, published by the xzf-thu team under Apache-2.0 and tracked on GitHub at 2.1k stars and 140 forks. Instead of dumping every exchange into one retrieval database, it splits memory across two coordinated stores: a left brain that organizes factual memories into schemas and entities for precise lookups, and a right brain that tracks persona, emotion, and relationships through standalone and cross-entity nodes. The project's stated pitch is that a voice assistant should remember not just what you said but who you are and how you felt saying it. The pipeline is streaming end to end. While you are still talking, VoiceMem runs audio segmentation, ASR, memory extraction, and structured writes into a memory graph; retrieval uses speculative prefetching (0–300 ms) so the query often resolves before the speaker finishes the sentence. On the numbers the project publishes, that yields a 134 ms response time against Mem0's 1,440 ms and roughly 300–430 memory tokens per query turn against 6,956 for Mem0 and 1,899 for EverMemOS. It is multi-modal by design: VoiceMem can remember voice, speaker identity, sound events, multi-party conversations, and music from real-world audio. The repo ships built-in modules for ASR, speaker verification, scene detection, emotion recognition, and local embedding, and every component — including the underlying memory engine and the TTS backend — is decoupled and swappable. Around it sit the VoiceMem Utils package for plug-and-play use, the VoiceMem model families (including a fine-tuned Qwen reply model) that read the memories it produces, and the ChatMem-400K dataset for fine-tuning. Accuracy claims are worth noting for buyers: the project reports 91.2% on LoCoMo with only Top-5 memories (Mem0 at 61.68%) and 69.44% on PersonaMem. VoiceMem targets developers and researchers building voice assistants that need persistent, emotionally aware memory at low latency and low token cost — not teams shopping for a hosted SaaS. Alternatives like Mem0 and Zep give you a managed service; VoiceMem gives you the source code and asks you to run it yourself.
Behind the Verdict
VoiceMem's core idea is the split between a left brain and a right brain. The left brain stores factual schemas and entities and, per the README, maintains Mem0's full-load performance under a Top-3 constraint. The right brain handles what the project calls 'EQ' — long- and short-term emotional attribution through cross-entity nodes that are maintained in conjunction with left-brain facts. For a voice agent, that distinction matters: a factual memory alone tells the agent you are a vegetarian with a nut allergy; the right-brain layer lets it notice you sounded frustrated the last time it suggested a restaurant. The practical win is token cost and latency. The README claims roughly 300 tokens per single-turn query and 0–300 ms speculative prefetching inside a voice turn, and the seed data cites 134 ms total response against Mem0's 1,440 ms. If you're paying per token on a high-volume voice product, the ~430 vs ~6,956 memory-token delta is not a rounding error. The architectural reason is that querying runs on pure vector retrieval and is decoupled from the slower write path (fact extraction, tagging, graph construction), so a slow ingest doesn't slow your response. Where VoiceMem is genuinely unusual is how much of it you can replace. The README states the whole architecture is decoupled and every component — including the underlying memory engine — can be swapped. Barge-in handling is two-stage: VAD pauses and preserves the audio queue, clearing on a stop signal or stable ASR text. Output uses a PCM-sample-based timeline with AudioWorklet render progress so interrupts land where you expect. There's also a SessionBuffer that isolates per-session context and purges temporary conversations at session end, plus a web demo (web/run.py on localhost:8787) and evaluation code for LoCoMo and PersonaMem. The honest counterpoints: this is v0.0.2 (released 09/01/2026, following v0.0.1 on 08/27/2026), documentation may be incomplete, and APIs can move. Performance numbers come from the project's own technical report, not independent replication — treat them as promising leads, not settled verdicts. You will download models, wait for lazy-loaded local models to warm up, and debug at source level; there is no managed cloud service or commercial support. If your project needs an SLA or a hosted API, look at Mem0 or Zep instead and come back to VoiceMem if token economics become the binding constraint.
Researching VoiceMem? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas VoiceMem actually fits — and what changes day-one when you adopt it.
You clone the repo, pip install voicemem, download the default models via hf download, then construct a VoiceMem instance in mode="normal" and call vm.warmup() before feeding your first audio.
Outcome: You ingest a sample WAV (vm.ingest) and query it (vm.search), seeing result_leftbrain and result_rightbrain come back separately within minutes of a first install.
You build against the streaming interface — vm.stream(src_rate=..., vad_threshold=0.5, on_partial=...) — feeding 32ms PCM chunks and watching for the partial transcript to cross your SPEC_MIN_CHARS threshold, which triggers background retrieval before the user finishes the sentence.
Outcome: Memory lookups resolve during the utterance, so the agent answers with ~134 ms response time and ~300 tokens per turn instead of waiting for end-of-speech.
You run the repo's evaluation code for LoCoMo and PersonaMem, and use the ChatMem-400K dataset plus finetune pipeline to train a VoiceMem model on your own conversational data.
Outcome: You get comparable benchmark numbers and a fine-tuned model family that reads the memories your pipeline produces.
Use Cases
- Build a voice assistant that remembers user preferences and past conversations over long periods.
- Create an empathetic voice agent that tracks emotional states and responds appropriately.
- Implement a low-latency memory layer for real-time voice interaction with speculative prefetching.
- Fine-tune VoiceMem language models on your own conversational data using the ChatMem-400K dataset.
- Replace the built-in memory engine with your own while keeping the dual-brain architecture intact.
- Integrate VoiceMem into open-source voice agent frameworks for personalized interactions.
Models Under the Hood
as of 2026-09-21
Limitations
- As of September 2026, VoiceMem is in early development (v0.0.2) and may have incomplete documentation or unstable APIs.
- The project's performance claims (e.g., low latency, token efficiency) come from the project's own README and technical report and may need independent verification.
- Being open-source infrastructure, it requires technical expertise to install, configure, and integrate into existing voice agent pipelines; there is no managed cloud service or commercial support.
- The current features and models are tailored to the provided dataset and may need fine-tuning for specific use cases.
as of 2026-09-22
Verification history
We have re-verified VoiceMem 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
Showing the 6 most recent of 7 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published VoiceMem tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Open Source (Apache-2.0)
$0
Ideal for
Python developers and research teams self-hosting voice-agent memory who can absorb the model downloads, warmup, and integration work
What this tier adds
Starting tier — the only tier: full memory system source, built-in ASR/voiceprint/scene/emotion/embedding modules, and optional finetuned Qwen reply model under Apache-2.0
Where the pricing makes sense
The company stage and team size where VoiceMem's pricing actually pencils out — and where peers do it cheaper.
VoiceMem is free under Apache-2.0, which puts it below managed alternatives like Mem0 and Zep (typically per-request or per-seat SaaS pricing) but shifts cost to your infrastructure and engineering hours. It's best when token economics are the binding constraint: at ~430 memory tokens per turn vs Mem0's ~6,956, a high-volume voice product can save real money on LLM calls. If you need vendor support and an SLA, a paid managed service is the right comparison.
Setup time & first value
How long it actually takes to get something useful out of VoiceMem — broken out by persona, not the marketing-page minute.
Python developers: roughly 15–30 minutes to first ingest (clone, pip install voicemem, hf download the default models, warm up). Researchers adding evaluation or finetuning: plan a half day for model downloads and eval runs. Teams wiring VoiceMem into an existing voice agent: budget a day or two for source-level integration and debugging, given v0.0.2 APIs.
Switching to or from VoiceMem
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From Mem0: map your existing fact memories into VoiceMem's left brain, and keep Mem0 running in parallel until the dual-brain writes stabilize.
- →From flat transcript retrieval: ingest your historical audio or transcripts through vm.ingest so the extraction pipeline builds schemas and entities.
- →From EverMemOS: port your retrieval calls to vm.search and use Top-K to control context length.
- →From a custom vector store: keep the vector layer but replace the write pipeline with VoiceMem's streaming extractor to get speculative prefetching.
- ↗To Mem0: swap the memory engine behind the architecture — VoiceMem's README states every component including the memory engine is replaceable.
- ↗To Zep: export the left-brain fact graph and reindex into Zep's managed store if you decide you need an SLA.
- ↗To your own engine: keep the ASR, speaker, emotion, and scene modules and swap only the memory store via the decoupled component interface.
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “VoiceMem”, and we withheld 6: 6 could not be judged, because “VoiceMem” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about VoiceMem.
Official links
Tools that pair well with VoiceMem
Common stack mates teams adopt alongside VoiceMem, with the specific reason each pairing earns its keep.
Featured Head-to-Head Comparisons
Voicemem vs Composio Mcp
These don't compete. Composio MCP is the plumbing that lets a coding agent (Claude, Cursor, Codex, ChatGPT) actually send the email, update the Salesforce record, or open the PR — priced freemium and climbing to $99/mo when your team outgrows the free tier. VoiceMem is an Apache-2.0 memory engine for people building real-time voice assistants who want their agent to remember not just facts but tone and relationships — free, self-hosted, and explicitly a v0.0.2 research project with no SLA. Buy Composio if your problem is app reach. Adopt VoiceMem if your problem is low-latency, emotionally aware recall in a voice loop. Nobody is choosing one over the other.
Voicemem vs Vectorize
These aren't really substitutes — they overlap only at the abstract level of 'agent memory.' Pick Vectorize if you're building text or MCP-based agents (Claude Code, Cursor, Google ADK) and want memory that turns failures into reusable judgment, with a free MIT self-host path and a managed cloud option when you'd rather not run infrastructure. Pick VoiceMem only if you're building a real-time voice agent that needs sub-150ms streaming memory with persona and emotion — and you're comfortable with a v0.0.2 research project, self-hosting, and benchmark claims that come from the vendor's own report.
Voicemem vs Mem0
These overlap less than the labels suggest. Mem0 is the buyer's pick if you want a managed, cross-session memory layer for text-based agents, chatbots, and CRMs today — you pay for reliability, K8s/air-gapped self-hosting, and a Pro tier at $249/mo. VoiceMem is the pick if you're building a real-time voice agent and care about per-turn token cost, emotional/persona memory, and low latency — but you're adopting a v0.0.2 research project with no SLA and self-hosting required. Don't pick VoiceMem expecting production support, and don't pick Mem0 expecting audio-native dual-brain retrieval.
Voicemem vs Arcade Ai
These aren't competitors — you would not swap one for the other, and a comparison page is the wrong place to decide. Arcade AI is a buy-an-enterprise-runtime decision for teams whose agents touch real user accounts and need per-action audit trails naming agent, user, and system. VoiceMem is a self-host, Apache-2.0 memory layer for people building a real-time voice agent who want persona and emotion recall at low latency, and who accept v0.0.2 maturity with no support. If you're asking "which of these do I pay for," the answer depends entirely on whether your problem is authorization or memory — and if it's both, you'd run them in different parts of your stack, not pick between them.
Voicemem vs Tobira
These aren't competitors — don't shortlist them against each other. Tobira is for people who want a public, censorship-resistant address and machine-readable profile for an AI agent (so other agents and crawlers can find and talk to it). VoiceMem is for builders wiring persistent, emotionally aware memory into a real-time voice agent. If you need agent discoverability, take Tobira. If you need voice memory with published latency and accuracy numbers, take VoiceMem. If you need both, use both.
Voicemem vs Granica Ai
These are not competitors — Granica sells an enterprise efficiency layer for petabyte-scale tabular data lakes and long-running agents, while VoiceMem is a free Apache-2.0 memory backend for real-time voice agents. Pick Granica if you are an enterprise data engineer trying to cut Iceberg/Delta Lake storage and processing spend (with a cloud perimeter and petabyte-scale data), or an AI team needing agent state persistence via Myelin. Pick VoiceMem if you are building a self-hosted voice agent prototype and need emotion-aware memory at low latency, and you can tolerate v0.0.2 research-grade code with no SLA.
Alternatives to VoiceMem
View allPopular in Agent Memory & Runtimes
Frequently Asked Questions
Categories
Best-of guides
Used VoiceMem? Help shape our editorial sentiment research.