VoiceMem vs Tobira

Side-by-side comparison of features, pricing, and ratings

Analysis reviewed Live tool data as of 2026-09-30
Cross-checked through our multi-step verification ·
Saved

At a glance

DimensionVoiceMemTobira
What it actually isOpen-source dual-brain memory for voice agentsIdentity/discovery network for AI agents
PricingFree (Apache-2.0, self-hosted)Free (permissionless registration)
Core primitiveLeft-brain facts / right-brain persona memory stores@handle + profile + JSON discovery files
Latency claim134 ms response (vs Mem0's 1,440 ms)Not a latency product
Benchmarks91.2% LoCoMo (Top-5), 69.44% PersonaMemNone published
Maturityv0.0.2 research projectLaunched 2026-06-28
VoiceMem
VoiceMem

Open-source dual-brain memory for real-time voice agents — facts in the left brain, emotion in the right, streaming at 134ms.

Visit Website
Tobira
Tobira

Tobira gives AI agents public addresses — @handles, public profiles, guest chat, and machine-readable discovery files like agent.json and llms.txt.

Visit Website
Pricing
Free
Free
Plans
$0
$0/mo
Popularity
2 views
5.7k views
Skill Level
Advanced
Beginner-friendly
API Available
Platforms
APICLIWeb
Web
Categories
🧠 Agent Memory & Runtimes
🧠 Agent Memory & Runtimes🔌 MCP Servers & Agent Tooling
Features
Dual-brain memory: left brain stores factual schemas and entities, right brain stores persona, emotion, relationships
Fully streaming pipeline: audio segmentation, ASR, memory extraction, and graph writes while the user speaks
Speculative prefetching inside a voice turn (0–300 ms) so retrieval starts before the user finishes
Top-K memory routing and ranking controls context length
~300 tokens per single-turn query (project benchmark ~430 memory tokens per turn)
Published 134 ms response time vs Mem0's 1,440 ms
91.2% on LoCoMo with Top-5 memories (Mem0: 61.68%) and 69.44% on PersonaMem
Multi-modal memory from real audio: voice, speaker, sound events, multi-party conversations, music
Built-in ASR, speaker verification, scene detection, emotion recognition, and local embedding modules
Swappable components including the underlying memory engine and a pluggable TTS layer
SessionBuffer isolates per-session context, with temporary conversations purged at session end
Two-stage barge-in: VAD pauses and preserves audio queue, clears on stop or stable ASR text
PCM-sample-based output timeline with AudioWorklet render progress for interrupt handling
VoiceMem official model families (fine-tuned Qwen reply model) that read WireMem memories
ChatMem-400K dataset plus finetune pipeline and evaluation scripts for custom training
Claim a public @handle as an AI agent address
Publish a profile with offers, needs, goals, blockers, and proof
Guest chat for human-to-agent conversation
agent.json machine-readable discovery file
guest-agent.json for visiting-AI JSON access
llms.txt for AI assistants
AGENTS.md discovery file for agents
Link headers and discovery metadata for agent lookup
Site Agent representing a website or business
Attached Knowledge for approved agent context
Channel instructions to route agent conversations
Permissionless registration on an open network
Profile pages as public address surfaces

Feature-by-feature

Tobira and VoiceMem sit at opposite ends of the agent stack. Tobira is a discovery and identity layer: you claim a public @handle, publish a profile (offers, needs, goals, blockers, proof), and expose machine-readable surfaces — agent.json, guest-agent.json, llms.txt, AGENTS.md, Link headers, discovery metadata — so other agents and crawlers can find and read you. It also ships Guest chat for human-to-agent conversation, Site Agents for websites and businesses, Attached Knowledge for approved context, and channel instructions to route conversations. Registration is permissionless and censorship-resistant.

VoiceMem is a memory runtime. Its distinguishing feature is a dual-brain split: a left brain holding factual schemas and entities, a right brain holding persona, emotion, and relationships. The pipeline is streaming end to end — audio segmentation, ASR, memory extraction, and graph writes happen while the user speaks — with speculative prefetching at 0–300 ms inside a turn. It supports Top-K memory routing (~430 memory tokens per query turn), multi-modal audio memory (voice, speaker, sound events, multi-party, music), built-in ASR/speaker verification/scene detection/emotion recognition, swappable memory engine and TTS, SessionBuffer isolation, and a two-stage barge-in design.

The overlap is effectively zero. Tobira answers "who is this agent and how do I reach it?"; VoiceMem answers "what does this voice agent remember about you, and how fast?". Neither substitutes for the other.

Pricing compared

Both are free, but free at very different price-of-admission levels. Tobira's model is permissionless, censorship-resistant registration — there's no paywall described, and the explicit positioning is against "paid quotas" or "guaranteed leads" right now. Cost of entry for a buyer is essentially your time: claim an @handle, publish a profile, wire up the JSON surfaces.

VoiceMem is Apache-2.0 open source, which means the license is free but the real bill is operational. You self-host; you pay in engineering hours, model downloads, local warmup, and source-level debugging. The docs are candid: this is v0.0.2 and there is no managed service, no vendor SLA, no hosted API. Token efficiency is part of the cost story — ~430 memory tokens per query turn via Top-K routing — which matters if you're paying per token upstream.

The comparison that counts on the VoiceMem side is against Mem0 (1,440 ms vs 134 ms response; 61.68% vs 91.2% on LoCoMo with Top-5). Those are vendor-reported numbers from an open technical report with eval scripts and ChatMem-400K attached, not independent third-party replication — so budget for your own validation. Neither tool has a paid tier to upgrade into.

Who should pick which

  • Web3 developer building decentralized agent identity
    Pick: Tobira

    Permissionless, censorship-resistant @handle registration plus agent.json/AGENTS.md discovery is exactly the identity substrate you'd otherwise have to build.

  • Site owner whose business keeps getting misread by visiting AI assistants
    Pick: Tobira

    A Site Agent plus llms.txt, guest-agent.json, and Attached Knowledge gives crawlers approved, machine-readable context about your business.

  • Engineer building a real-time voice agent with persistent memory
    Pick: VoiceMem

    Streaming extraction, 134 ms response, dual-brain fact/persona split, and Top-K routing are purpose-built for low-latency voice turns.

  • Researcher studying voice AI memory architectures
    Pick: VoiceMem

    The open technical report, eval scripts, ChatMem-400K, and swappable components make it reproducible rather than a black box.

  • Team that needs production SLA, support, and a managed API today
    Pick: Tobira

    VoiceMem explicitly excludes these (v0.0.2 research project, no hosted service). Tobira's free network at least gives you something shipping now — though neither is a turnkey managed product.

Frequently Asked Questions

Could I replace one of these with the other?

No. Tobira hands out addresses and profiles so agents can be found; VoiceMem stores what a voice agent remembers mid-conversation. There is no feature overlap to migrate.

Do I need Tobira to use VoiceMem?

Only if you want your voice agent to be publicly discoverable by other agents or crawlers. VoiceMem works entirely as a self-hosted memory backend without any public identity layer.

Are the VoiceMem benchmarks trustworthy?

They come from the vendor's own technical report, and the project's own 'not for' list flags teams requiring independent third-party replication. Treat 134 ms, 91.2% LoCoMo, and 69.44% PersonaMem as claims to reproduce with the supplied eval scripts.

What does Tobira cost if I scale up?

Nothing described. Registration is permissionless and there's no paid quota model in the current data. Your cost is the engineering effort to publish profiles and JSON surfaces.

Which is safer to bet on for a deadline next month?

Tobira shipped publicly in June 2026 and is usable immediately on its free tier. VoiceMem explicitly warns against deadlines that leave no room for model downloads, local warmup, and source-level debugging at v0.0.2.

Can VoiceMem's components be swapped out?

Yes — the memory engine itself and the TTS backend are both pluggable, and SessionBuffer isolates per-session context with temporary conversations purged at session end.

What does Tobira expose to other machines?

agent.json, guest-agent.json, llms.txt, AGENTS.md, plus Link headers and discovery metadata — multiple machine-readable surfaces so agents and crawlers can locate and read an @handle.

More VoiceMem or Tobira comparisons

Explore each tool further

Browse these categories

Still deciding? Get the weekly AI tools brief

One email a week — new tools, honest comparisons, no spam.

Last reviewed: September 21, 2026