Vocalinux

Vocalinux

Free, AGPL-3.0 system-tray voice dictation for Linux desktops, with local speech engines and no account required

69/100MonitorFreeFree

If you type on Linux and care where your audio goes, install this. Local engines are the default, whisper.cpp gets Vulkan acceleration on AMD and Intel as well as NVIDIA, and v0.17.0 added Faster Whisper (CTranslate2/INT8, CPU) plus Parakeet TDT 0.6B so modest hardware has a path too. The tradeoff is scope: Linux desktops only — VocaHQ ships separate apps for other platforms — and no cloud-style extras like smart replies or meeting summaries. Against cloud dictation subscriptions it wins on privacy and cost; against a hand-rolled whisper.cpp hotkey script it wins on tray, settings and reliable injection across X11 and Wayland. For hands-free dictation in terminals and editors, it is the

Verified 5d ago · liveness 69/100 · cite: rightaichoice.com/tools/vocalinux

Best for
  • Linux desktop users who want private, offline voice dictation with no subscription
  • Developers dictating into terminals, IDEs and browsers without leaving the keyboard
  • AMD or Intel GPU owners who want GPU-accelerated transcription, not just CUDA
  • Users with RSI or accessibility needs who need system-wide hands-free text input
Not ideal for
  • Windows or macOS users — VocaHQ ships separate VocaMac and Windows beta apps, not this one
  • People who want a double-click install with no terminal involvement
  • Anyone depending on cloud dictation extras like smart replies, translation or meeting summaries
Visit Website

IntermediateRoughly 1–2 minutes for the default whisper.cpp setup, since the installer detects your hardware and the Tiny model is about 74MB. Budget longer if you pick a larger model, or add Faster Whisper or Parakeet as extras via installer --engine flags. AppImage users need host text-injection tools installed first; Snap on native Wayland needs the uinput plug connected.Desktop · CLIAPI availableVerified 5d ago
Pricing
Free
FreeFree tier
Learning curve
Intermediate
Roughly 1–2 minutes for the default whisper.cpp setup, since the installer detects your hardware and the Tiny model is about 74MB. Budget longer if you pick a larger model, or add Faster Whisper or Parakeet as extras via installer --engine flags. AppImage users need host text-injection tools installed first; Snap on native Wayland needs the uinput plug connected.
Runs on
DesktopCLI
API available
Who it's for
Linux developer on an AMD GPUPrivacy-conscious writer on a modest laptopAccessibility user on KDE Plasma Wayland
Live sentiment
Is Vocalinux actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Vocalinux if you don't run a Linux desktop, or if you want cloud dictation extras like translation, smart replies or meeting summaries rather than on-device transcription into the focused window.

The 30-second take
Price reality

Vocalinux is free and open source under AGPL-3.0, with the full app and every local engine included — whisper.cpp, Faster Whisper, Whisper, VOSK, Parakeet and Remote API. The costs it does carry are indirect: disk space for downloaded models (Tiny ships around 74MB, larger variants more), your own GPU or CPU time for local transcription, and any server you run yourself for Remote API. That puts it below paid cloud dictation subscriptions and roughly on par with hand-rolling whisper.cpp, minus

In short

Vocalinux — Free, AGPL-3.0 system-tray voice dictation for Linux desktops, with local speech engines and no account required. Best for Linux desktop users who want private, offline voice dictation with no subscription, Developers dictating into terminals, IDEs and browsers without leaving the keyboard, AMD or Intel GPU owners who want GPU-accelerated transcription, not just CUDA. Free to use.

What's new in Vocalinux

Checked 5 days ago

Across the latest 5 updates: 3 feature updates, 1 launch and 1 changelog entry.

What people actually say about Vocalinux — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

36 mentions across 4 sources (Hacker News, YouTube, Bluesky, GitHub) · researched Jul 6, 2026.

49% positive51% critical

Average across the 4 sources that answered — each source counts once, not each post.

Recurring strengths
  • +100% offline and private — no data ever leaves your machine.
  • +Supports multiple STT engines: whisper.cpp, VOSK, and OpenAI Whisper.
  • +GPU acceleration via Vulkan works on AMD, Intel, and NVIDIA GPUs.
  • +Compatible with both X11 and Wayland display servers.
  • +One-command installer supports Ubuntu, Fedora, Debian, Arch, and openSUSE.
Recurring frustrations
  • −Frequent ghost/hallucination text after silence (GitHub issue reported).
  • −Non-US keyboard layouts cause incorrect character injection with ydotool.
  • −Installation often completes but the app fails to launch or respond.
  • −Audio recording may start but never generate transcription (common bug).
  • −38 open GitHub issues — many core functionality bugs unresolved.
Patterns worth knowing
Installation and launch failures plague new users
Seen on GitHub, Bluesky
Hallucination/ghost text is a major reliability concern
Seen on GitHub
Non-US keyboard layouts render dictation unusable
Seen on GitHub
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • No hidden costs, but requires time to troubleshoot bugs

Viability Score

69/100
Monitor

How well maintained and how widely used is Vocalinux? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
49
What the vendor publishes
20

Last calculated: October 2026

How we score →

Key Features

  • System-wide offline voice dictation into any focused app
  • Hold Right Alt push-to-talk by default; toggle mode available
  • whisper.cpp engine with Vulkan GPU acceleration on AMD, Intel and NVIDIA
  • Default whisper.cpp Tiny model at roughly 74MB
  • Original OpenAI Whisper via PyTorch/CUDA for NVIDIA workflows
  • VOSK engine with ~40MB models for older, low-RAM machines
  • Faster Whisper (CTranslate2/INT8 CPU) opt-in engine, installer --engine=faster_whisper
  • Parakeet TDT 0.6B via sherpa-onnx, installer --engine=parakeet
  • Remote API engine targeting OpenAI-compatible or whisper.cpp servers
  • Silero VAD neural silence filtering before transcription
  • X11 and Wayland support with IBus injection and clipboard fallback
  • ~33 languages with searchable list, auto-capitalize and locale-seeded first run
  • Localized voice punctuation commands for it/fr/de/es/pt/nl/pl/ru
  • Bare F1–F24 push-to-talk shortcuts
  • Snap packaging with ydotool and uinput for native Wayland

About Vocalinux

FreeIntermediateAPI availableDesktop · CLI

Vocalinux is a free, open-source (AGPL-3.0) system tray app that adds system-wide voice dictation to Linux desktops. You hold Right Alt (or switch to toggle mode), speak, and the transcript lands in whichever window already has focus — a terminal, browser, IDE, or office document — on both X11 and Wayland. Audio is processed by local speech engines on your own machine. whisper.cpp is the default engine (C++ Whisper with Vulkan acceleration, so AMD, Intel, and NVIDIA GPUs are all covered), and the default Tiny model is roughly 74MB. Other local engines include the original OpenAI Whisper on PyTorch/CUDA for NVIDIA workflows, VOSK with ~40MB models for older or low-RAM machines, Parakeet, and — as of v0.17.0 — Faster Whisper (CTranslate2/INT8, CPU) and Parakeet TDT 0.6B via sherpa-onnx, both opt-in via installer --engine flags. If you would rather offload compute, Remote API points at an OpenAI-compatible endpoint or a whisper.cpp server you configure while desktop text injection stays local. Day-to-day use runs through a tray icon: hold-to-record or toggle mode, configurable hotkeys (bare F1–F24 shortcuts as of v0.17.0), a searchable settings dialog, a searchable language list, and an in-app update checker. Silero VAD filters silence before transcription, auto-capitalize cleans up output, and the language catalog spans roughly 33 languages. Voice punctuation commands are now localized for Italian, French, German, Spanish, Portuguese, Dutch, Polish and Russian. Injection uses IBus where available, with clipboard fallback on unbridged compositors. Installation is one interactive command that detects your hardware and wires up the desktop app; AppImage (x86_64 and aarch64), Flatpak, Snap and AUR packages are alternatives. Ubuntu, Fedora, Debian, Arch and openSUSE are the supported targets. It is desktop-first and Linux-only: unlike cloud dictation services there is no required account and the installed app sends no usage telemetry, and unlike skeletal hotkey scripts it ships a real tray, settings dialog and engine picker.

Behind the Verdict

Vocalinux solves a narrow problem well: getting words into the field that already has focus on a Linux desktop without a cloud round-trip. The architecture is the selling point. Microphone audio is held in memory for the recording, transcribed by a local engine, and injected at the cursor; a Remote API is a separate, optional stop that only ever talks to a server you configure. The installed app sends no usage telemetry, and the whole injection path is inspectable because the project is AGPL-3.0 on GitHub. Engine choice is genuinely useful rather than decorative. whisper.cpp is the default and covers AMD, Intel and NVIDIA GPUs through Vulkan, which matters if you are not on CUDA. OpenAI Whisper via PyTorch/CUDA is there for people already in the CUDA stack. VOSK trades accuracy for a ~40MB footprint on older machines. v0.17.0 added two more: Faster Whisper on CTranslate2/INT8 for CPU, and Parakeet TDT 0.6B via sherpa-onnx, both installed with --engine flags. Remote API round-trips to an OpenAI-compatible endpoint or a whisper.cpp server you run. The desktop integration is where most Linux dictation attempts fall over, and this one has been iterating on exactly that. Injection supports X11 and Wayland, uses IBus where available and a clipboard fallback on unbridged compositors. Recent releases fixed KDE cases where leftover IBus was not the session input method, made wtype and ydotool deliver real chord shortcuts on Wayland, added layout-aware Ctrl+V paste and XWayland clipboard paste instead of layout-garbled xdotool type, and gave terminals Ctrl+Shift+V. v0.17.0 added Snap packaging that bundles ydotool and uinput for native Wayland — with the caveat that the GitHub .snap sideload exists while Store review of the uinput plug is pending, and Snap 0.16.2 edge has no uinput plug at all, so dictation there only reaches XWayland apps. The weaknesses are scope and setup. It is Linux-only: the same family ships VocaMac and Windows apps separately, so do not look here for either. Installation is fine from the interactive installer but it is still a terminal command, and host text-injection tools are required if you go the AppImage route. Language behaviour depends on the engine and model you pick — v0.17.0 changed the UI so language and speed/accuracy come first and engine/size sit under Advanced, and it fixed a case where English-only models hid other languages; choosing Polish or auto-detect now switches to the multilingual sibling. Local engines also mean local compute: on very old hardware, even VOSK can be too much. Where it fits: developers dictating into terminals and IDEs, note-taking into any focused app, RSI and accessibility users who need system-wide hands-free input, and privacy-conscious teams that cannot send audio to a cloud speech service. Where it does not: anyone wanting a double-click GUI install, mobile or browser dictation, or cloud extras like translation and meeting summaries — none of that is here.

Researching Vocalinux? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Vocalinux actually fits — and what changes day-one when you adopt it.

Linux developer on an AMD GPU

Installs Vocalinux with the interactive installer, keeps whisper.cpp with its Vulkan backend and the Tiny model, and holds Right Alt while dictating code comments into the IDE that already has focus.

Outcome: Hands-free comments and docstrings land at the cursor without leaving the keyboard or sending audio off the machine.

Privacy-conscious writer on a modest laptop

Installs the VOSK engine with its ~40MB models to keep RAM use low, or runs Faster Whisper on CTranslate2/INT8 with the CPU, then uses toggle mode to dictate reports into an office document.

Outcome: Dictation works on hardware without a strong GPU, and nothing is uploaded to a speech cloud.

Accessibility user on KDE Plasma Wayland

Uses the tray icon and a bare F-key push-to-talk shortcut to dictate into Kate, browsers and terminals after the 0.16.2 KDE and Wayland shortcut fixes.

Outcome: System-wide voice input replaces typing for most text entry, with IBus injection where available and clipboard fallback where it isn't.

Use Cases

Models Under the Hood

whisper.cppOpenAI WhisperFaster WhisperVOSKParakeetParakeet TDT 0.6B

as of 2026-09-23

Limitations

  • Vocalinux is a Linux-only desktop application for X11 and Wayland, installed via a command-line installer with Flatpak, AppImage, Snap and AUR packages as alternatives.
  • Its focus is offline, on-device dictation with an optional Remote API pointing at an OpenAI-compatible or whisper.cpp server.
  • Language support depends on the selected engine and model; changelog entries note that choosing some languages switches to multilingual model variants, and that English-only models previously hid other languages (fixed in v0.17.0).
  • Snap on the 0.16.2 edge channel ships without the uinput plug, so dictation there only reaches XWayland apps until a later store snap with ydotool is installed and uinput is connected.
  • Faster Whisper and Parakeet are opt-in extras rather than defaults.

as of 2026-10-03

Verification history

We have re-verified Vocalinux 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-checked, vendor evidence unchanged

Showing the 6 most recent of 8 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Vocalinux tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Open Source (AGPL-3.0)

$0

Ideal for

Individual Linux desktop users and small teams who want offline dictation with no account and no per-seat cost.

What this tier adds

Starting tier: the full app at $0, including every local engine (whisper.cpp, Faster Whisper, Whisper, VOSK, Parakeet) and Remote API support.

Where the pricing makes sense

The company stage and team size where Vocalinux's pricing actually pencils out — and where peers do it cheaper.

Vocalinux is free and open source under AGPL-3.0, with the full app and every local engine included — whisper.cpp, Faster Whisper, Whisper, VOSK, Parakeet and Remote API. The costs it does carry are indirect: disk space for downloaded models (Tiny ships around 74MB, larger variants more), your own GPU or CPU time for local transcription, and any server you run yourself for Remote API. That puts it below paid cloud dictation subscriptions and roughly on par with hand-rolling whisper.cpp, minus

Setup time & first value

How long it actually takes to get something useful out of Vocalinux — broken out by persona, not the marketing-page minute.

Roughly 1–2 minutes for the default whisper.cpp setup, since the installer detects your hardware and the Tiny model is about 74MB. Budget longer if you pick a larger model, or add Faster Whisper or Parakeet as extras via installer --engine flags. AppImage users need host text-injection tools installed first; Snap on native Wayland needs the uinput plug connected.

Switching to or from Vocalinux

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From hand-rolled whisper.cpp hotkey scripts: install Vocalinux and point it at the same engine family, gaining a tray, settings dialog and X11/Wayland injection.
  • →From cloud dictation subscriptions: install locally, pick an engine and model, and your audio stops leaving the machine for routine dictation.
  • →From a whisper.cpp server setup: keep the server and configure Remote API to send audio there while desktop injection stays local.
  • →From X11-only dictation tools: install and use IBus injection or the clipboard fallback to keep working under Wayland.
Migrating out
  • ↗To VocaMac or a Windows dictation app: Vocalinux is Linux-only, so other desktops need a separate product from the same family.
  • ↗To a cloud dictation service: if you need translation, smart replies or meeting summaries, export your notes and move to a service that provides them.
  • ↗To a manual whisper.cpp CLI workflow: uninstall with uninstall.sh and keep your model files for command-line transcription.

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Vocalinux”, and we withheld 6: 6 could not be judged, because “Vocalinux” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Vocalinux.

Tools that pair well with Vocalinux

Common stack mates teams adopt alongside Vocalinux, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Vocalinux

View all
Voicebox

Voicebox

Open-source desktop voice studio for local cloning, dictation, and agent speech — no account, no cloud, MIT licensed.

FreemiumTry
Hyprwhspr

Hyprwhspr

Hyprwhspr is free, MIT-licensed speech-to-text for Linux that runs local AI dictation models on Wayland and X11.

FreeTry
NexTalk

NexTalk

Open-source, 100% offline voice dictation for Linux desktops, integrated natively with Fcitx5.

FreeTry

Frequently Asked Questions

Used Vocalinux? Help shape our editorial sentiment research.