Vocalinux
Free, AGPL-3.0 system-tray voice dictation for Linux desktops, with local speech engines and no account required
If you type on Linux and care where your audio goes, install this. Local engines are the default, whisper.cpp gets Vulkan acceleration on AMD and Intel as well as NVIDIA, and v0.17.0 added Faster Whisper (CTranslate2/INT8, CPU) plus Parakeet TDT 0.6B so modest hardware has a path too. The tradeoff is scope: Linux desktops only — VocaHQ ships separate apps for other platforms — and no cloud-style extras like smart replies or meeting summaries. Against cloud dictation subscriptions it wins on privacy and cost; against a hand-rolled whisper.cpp hotkey script it wins on tray, settings and reliable injection across X11 and Wayland. For hands-free dictation in terminals and editors, it is the
Verified 5d ago · liveness 69/100 · cite: rightaichoice.com/tools/vocalinux
- Linux desktop users who want private, offline voice dictation with no subscription
- Developers dictating into terminals, IDEs and browsers without leaving the keyboard
- AMD or Intel GPU owners who want GPU-accelerated transcription, not just CUDA
- Users with RSI or accessibility needs who need system-wide hands-free text input
- Windows or macOS users — VocaHQ ships separate VocaMac and Windows beta apps, not this one
- People who want a double-click install with no terminal involvement
- Anyone depending on cloud dictation extras like smart replies, translation or meeting summaries
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Vocalinux if you don't run a Linux desktop, or if you want cloud dictation extras like translation, smart replies or meeting summaries rather than on-device transcription into the focused window.
Vocalinux is free and open source under AGPL-3.0, with the full app and every local engine included — whisper.cpp, Faster Whisper, Whisper, VOSK, Parakeet and Remote API. The costs it does carry are indirect: disk space for downloaded models (Tiny ships around 74MB, larger variants more), your own GPU or CPU time for local transcription, and any server you run yourself for Remote API. That puts it below paid cloud dictation subscriptions and roughly on par with hand-rolling whisper.cpp, minus
In short
Vocalinux — Free, AGPL-3.0 system-tray voice dictation for Linux desktops, with local speech engines and no account required. Best for Linux desktop users who want private, offline voice dictation with no subscription, Developers dictating into terminals, IDEs and browsers without leaving the keyboard, AMD or Intel GPU owners who want GPU-accelerated transcription, not just CUDA. Free to use.
What's new in Vocalinux
Checked 5 days agoAcross the latest 5 updates: 3 feature updates, 1 launch and 1 changelog entry.
Vocalinux v0.17.0 adds Faster Whisper and Parakeet TDT engines
Faster Whisper (CTranslate2/INT8, CPU) and Parakeet TDT 0.6B via sherpa-onnx ship as opt-in engines, installed with installer --engine flags. The GPU backend no longer installs the CUDA toolkit.
Vocalinux v0.17.0 reworks Speech Model setup and language picker
Setup now shows language and speed/accuracy first with engine and size under Advanced; the language picker filters while open and first run seeds the recognition language from your keyboard layout or locale.
Vocalinux v0.17.0 adds localized punctuation commands and F1–F24 shortcuts
Punctuation voice commands are localized for Italian, French, German, Spanish, Portuguese, Dutch, Polish and Russian, and bare F1–F24 keys can be bound as push-to-talk shortcuts.
Vocalinux v0.17.0 ships Snap packaging for native Wayland
A Snap build bundles ydotool and uinput for native Wayland dictation, distributed as a GitHub .snap sideload while Store review of the uinput plug is pending. Flatpak artifacts are now attached to releases too.
Vocalinux v0.16.2 fixes KDE, Wayland and IBus shortcuts
Leftover IBus is skipped when it is not the session input method so dictation types into Kate, browsers and terminals; wtype and ydotool deliver real chord shortcuts; 'delete that' now sends real BackSpace key events.
What people actually say about Vocalinux — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
36 mentions across 4 sources (Hacker News, YouTube, Bluesky, GitHub) · researched Jul 6, 2026.
Average across the 4 sources that answered — each source counts once, not each post.
- +100% offline and private — no data ever leaves your machine.
- +Supports multiple STT engines: whisper.cpp, VOSK, and OpenAI Whisper.
- +GPU acceleration via Vulkan works on AMD, Intel, and NVIDIA GPUs.
- +Compatible with both X11 and Wayland display servers.
- +One-command installer supports Ubuntu, Fedora, Debian, Arch, and openSUSE.
- −Frequent ghost/hallucination text after silence (GitHub issue reported).
- −Non-US keyboard layouts cause incorrect character injection with ydotool.
- −Installation often completes but the app fails to launch or respond.
- −Audio recording may start but never generate transcription (common bug).
- −38 open GitHub issues — many core functionality bugs unresolved.
- • No hidden costs, but requires time to troubleshoot bugs
Viability Score
How well maintained and how widely used is Vocalinux? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- System-wide offline voice dictation into any focused app
- Hold Right Alt push-to-talk by default; toggle mode available
- whisper.cpp engine with Vulkan GPU acceleration on AMD, Intel and NVIDIA
- Default whisper.cpp Tiny model at roughly 74MB
- Original OpenAI Whisper via PyTorch/CUDA for NVIDIA workflows
- VOSK engine with ~40MB models for older, low-RAM machines
- Faster Whisper (CTranslate2/INT8 CPU) opt-in engine, installer --engine=faster_whisper
- Parakeet TDT 0.6B via sherpa-onnx, installer --engine=parakeet
- Remote API engine targeting OpenAI-compatible or whisper.cpp servers
- Silero VAD neural silence filtering before transcription
- X11 and Wayland support with IBus injection and clipboard fallback
- ~33 languages with searchable list, auto-capitalize and locale-seeded first run
- Localized voice punctuation commands for it/fr/de/es/pt/nl/pl/ru
- Bare F1–F24 push-to-talk shortcuts
- Snap packaging with ydotool and uinput for native Wayland
About Vocalinux
Vocalinux is a free, open-source (AGPL-3.0) system tray app that adds system-wide voice dictation to Linux desktops. You hold Right Alt (or switch to toggle mode), speak, and the transcript lands in whichever window already has focus — a terminal, browser, IDE, or office document — on both X11 and Wayland. Audio is processed by local speech engines on your own machine. whisper.cpp is the default engine (C++ Whisper with Vulkan acceleration, so AMD, Intel, and NVIDIA GPUs are all covered), and the default Tiny model is roughly 74MB. Other local engines include the original OpenAI Whisper on PyTorch/CUDA for NVIDIA workflows, VOSK with ~40MB models for older or low-RAM machines, Parakeet, and — as of v0.17.0 — Faster Whisper (CTranslate2/INT8, CPU) and Parakeet TDT 0.6B via sherpa-onnx, both opt-in via installer --engine flags. If you would rather offload compute, Remote API points at an OpenAI-compatible endpoint or a whisper.cpp server you configure while desktop text injection stays local. Day-to-day use runs through a tray icon: hold-to-record or toggle mode, configurable hotkeys (bare F1–F24 shortcuts as of v0.17.0), a searchable settings dialog, a searchable language list, and an in-app update checker. Silero VAD filters silence before transcription, auto-capitalize cleans up output, and the language catalog spans roughly 33 languages. Voice punctuation commands are now localized for Italian, French, German, Spanish, Portuguese, Dutch, Polish and Russian. Injection uses IBus where available, with clipboard fallback on unbridged compositors. Installation is one interactive command that detects your hardware and wires up the desktop app; AppImage (x86_64 and aarch64), Flatpak, Snap and AUR packages are alternatives. Ubuntu, Fedora, Debian, Arch and openSUSE are the supported targets. It is desktop-first and Linux-only: unlike cloud dictation services there is no required account and the installed app sends no usage telemetry, and unlike skeletal hotkey scripts it ships a real tray, settings dialog and engine picker.
Behind the Verdict
Vocalinux solves a narrow problem well: getting words into the field that already has focus on a Linux desktop without a cloud round-trip. The architecture is the selling point. Microphone audio is held in memory for the recording, transcribed by a local engine, and injected at the cursor; a Remote API is a separate, optional stop that only ever talks to a server you configure. The installed app sends no usage telemetry, and the whole injection path is inspectable because the project is AGPL-3.0 on GitHub. Engine choice is genuinely useful rather than decorative. whisper.cpp is the default and covers AMD, Intel and NVIDIA GPUs through Vulkan, which matters if you are not on CUDA. OpenAI Whisper via PyTorch/CUDA is there for people already in the CUDA stack. VOSK trades accuracy for a ~40MB footprint on older machines. v0.17.0 added two more: Faster Whisper on CTranslate2/INT8 for CPU, and Parakeet TDT 0.6B via sherpa-onnx, both installed with --engine flags. Remote API round-trips to an OpenAI-compatible endpoint or a whisper.cpp server you run. The desktop integration is where most Linux dictation attempts fall over, and this one has been iterating on exactly that. Injection supports X11 and Wayland, uses IBus where available and a clipboard fallback on unbridged compositors. Recent releases fixed KDE cases where leftover IBus was not the session input method, made wtype and ydotool deliver real chord shortcuts on Wayland, added layout-aware Ctrl+V paste and XWayland clipboard paste instead of layout-garbled xdotool type, and gave terminals Ctrl+Shift+V. v0.17.0 added Snap packaging that bundles ydotool and uinput for native Wayland — with the caveat that the GitHub .snap sideload exists while Store review of the uinput plug is pending, and Snap 0.16.2 edge has no uinput plug at all, so dictation there only reaches XWayland apps. The weaknesses are scope and setup. It is Linux-only: the same family ships VocaMac and Windows apps separately, so do not look here for either. Installation is fine from the interactive installer but it is still a terminal command, and host text-injection tools are required if you go the AppImage route. Language behaviour depends on the engine and model you pick — v0.17.0 changed the UI so language and speed/accuracy come first and engine/size sit under Advanced, and it fixed a case where English-only models hid other languages; choosing Polish or auto-detect now switches to the multilingual sibling. Local engines also mean local compute: on very old hardware, even VOSK can be too much. Where it fits: developers dictating into terminals and IDEs, note-taking into any focused app, RSI and accessibility users who need system-wide hands-free input, and privacy-conscious teams that cannot send audio to a cloud speech service. Where it does not: anyone wanting a double-click GUI install, mobile or browser dictation, or cloud extras like translation and meeting summaries — none of that is here.
Researching Vocalinux? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Vocalinux actually fits — and what changes day-one when you adopt it.
Installs Vocalinux with the interactive installer, keeps whisper.cpp with its Vulkan backend and the Tiny model, and holds Right Alt while dictating code comments into the IDE that already has focus.
Outcome: Hands-free comments and docstrings land at the cursor without leaving the keyboard or sending audio off the machine.
Installs the VOSK engine with its ~40MB models to keep RAM use low, or runs Faster Whisper on CTranslate2/INT8 with the CPU, then uses toggle mode to dictate reports into an office document.
Outcome: Dictation works on hardware without a strong GPU, and nothing is uploaded to a speech cloud.
Uses the tray icon and a bare F-key push-to-talk shortcut to dictate into Kate, browsers and terminals after the 0.16.2 KDE and Wayland shortcut fixes.
Outcome: System-wide voice input replaces typing for most text entry, with IBus injection where available and clipboard fallback where it isn't.
Use Cases
- Dictate code comments and documentation into your IDE hands-free.
- Compose emails and reports without typing to reduce RSI strain.
- Transcribe meeting notes directly into a text editor.
- Control text input in any application while multitasking.
- Enable accessible computing for users with limited mobility.
- Use voice input in terminals and office applications on Linux.
- On a KDE Plasma Wayland desktop, dictate into Kate, browsers and terminals.
- Offload transcription to your own whisper.cpp server over a trusted LAN with Remote API.
Models Under the Hood
as of 2026-09-23
Limitations
- Vocalinux is a Linux-only desktop application for X11 and Wayland, installed via a command-line installer with Flatpak, AppImage, Snap and AUR packages as alternatives.
- Its focus is offline, on-device dictation with an optional Remote API pointing at an OpenAI-compatible or whisper.cpp server.
- Language support depends on the selected engine and model; changelog entries note that choosing some languages switches to multilingual model variants, and that English-only models previously hid other languages (fixed in v0.17.0).
- Snap on the 0.16.2 edge channel ships without the uinput plug, so dictation there only reaches XWayland apps until a later store snap with ydotool is installed and uinput is connected.
- Faster Whisper and Parakeet are opt-in extras rather than defaults.
as of 2026-10-03
Verification history
We have re-verified Vocalinux 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Vocalinux tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Open Source (AGPL-3.0)
$0
Ideal for
Individual Linux desktop users and small teams who want offline dictation with no account and no per-seat cost.
What this tier adds
Starting tier: the full app at $0, including every local engine (whisper.cpp, Faster Whisper, Whisper, VOSK, Parakeet) and Remote API support.
Where the pricing makes sense
The company stage and team size where Vocalinux's pricing actually pencils out — and where peers do it cheaper.
Vocalinux is free and open source under AGPL-3.0, with the full app and every local engine included — whisper.cpp, Faster Whisper, Whisper, VOSK, Parakeet and Remote API. The costs it does carry are indirect: disk space for downloaded models (Tiny ships around 74MB, larger variants more), your own GPU or CPU time for local transcription, and any server you run yourself for Remote API. That puts it below paid cloud dictation subscriptions and roughly on par with hand-rolling whisper.cpp, minus
Setup time & first value
How long it actually takes to get something useful out of Vocalinux — broken out by persona, not the marketing-page minute.
Roughly 1–2 minutes for the default whisper.cpp setup, since the installer detects your hardware and the Tiny model is about 74MB. Budget longer if you pick a larger model, or add Faster Whisper or Parakeet as extras via installer --engine flags. AppImage users need host text-injection tools installed first; Snap on native Wayland needs the uinput plug connected.
Switching to or from Vocalinux
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From hand-rolled whisper.cpp hotkey scripts: install Vocalinux and point it at the same engine family, gaining a tray, settings dialog and X11/Wayland injection.
- →From cloud dictation subscriptions: install locally, pick an engine and model, and your audio stops leaving the machine for routine dictation.
- →From a whisper.cpp server setup: keep the server and configure Remote API to send audio there while desktop injection stays local.
- →From X11-only dictation tools: install and use IBus injection or the clipboard fallback to keep working under Wayland.
- ↗To VocaMac or a Windows dictation app: Vocalinux is Linux-only, so other desktops need a separate product from the same family.
- ↗To a cloud dictation service: if you need translation, smart replies or meeting summaries, export your notes and move to a service that provides them.
- ↗To a manual whisper.cpp CLI workflow: uninstall with uninstall.sh and keep your model files for command-line transcription.
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Vocalinux”, and we withheld 6: 6 could not be judged, because “Vocalinux” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Vocalinux.
Official links
Tools that pair well with Vocalinux
Common stack mates teams adopt alongside Vocalinux, with the specific reason each pairing earns its keep.
Voicebox
Open-source desktop voice studio for local cloning, dictation, and agent speech — no account, no cloud, MIT licensed.
Hyprwhspr
Hyprwhspr is free, MIT-licensed speech-to-text for Linux that runs local AI dictation models on Wayland and X11.
NexTalk
Open-source, 100% offline voice dictation for Linux desktops, integrated natively with Fcitx5.
Featured Head-to-Head Comparisons
Vocalinux vs Poke Interaction Co
Vocalinux and Poke serve completely different needs. Vocalinux is the right choice if you're a Linux user who wants private, offline voice dictation into any app—free and open-source. Poke is ideal if you want an AI assistant that lives in your messaging apps and manages email, calendar, tasks, and health data, with paid tiers for advanced automation. Your pick depends on whether you need pure voice input or a full personal assistant.
Vocalinux vs Guesty
Vocalinux and Guesty serve completely different needs. Vocalinux is a free, open-source voice dictation tool for Linux users requiring privacy and offline operation. Guesty is a paid property management platform for vacation rental hosts needing AI-driven automation. Choose based on your task: dictation vs. hospitality management.
Vocalinux vs Gem
Vocalinux and Gem serve entirely different purposes: one is a free, offline voice dictation tool for Linux, the other is a paid, AI-powered recruiting platform. Choose Vocalinux if you value privacy and need local speech-to-text on Linux; choose Gem if you're a recruiter seeking an all-in-one ATS/CRM with AI automation.
Alternatives to Vocalinux
View allFrequently Asked Questions
Categories
Best-of guides
Topics
Used Vocalinux? Help shape our editorial sentiment research.