Push-to-Talk Dictation on Linux With nerd-dictation
Sold alongside its own clip-on microphone, Vibe Coding Pad provides three mechanical controls in a compact macro pad. nerd-dictation's begin, end and cancel commands can turn those controls into a local Linux dictation workflow.
How does nerd-dictation begin and end work?
The project uses Vosk for offline speech recognition and an external audio capture tool. The begin command starts recording and recognition; end finishes the session and outputs the text; cancel discards it. The project explicitly recommends binding those commands to accessible shortcuts.
Its default approach starts a process for manual dictation, while suspend and resume can keep data in memory on slower systems. Choose one lifecycle and keep the model path, audio input and output method in a documented configuration.
How do you bind it to a held hardware key?
Use a Linux input layer that can run one reviewed user command on key down and another on key up. Down invokes begin once; up invokes end once. Add a lock so key repeat cannot start several processes. Bind the second pad key to cancel and make that route work even when the target application is not focused.
Test by directing output to standard output before allowing simulated input. This confirms capture and recognition without typing into an unknown window.
How can text land in a Wayland terminal?
nerd-dictation documents several output tools. Its X11 default uses xdotool, while ydotool, dotool and wtype provide documented paths for Wayland or broader Linux input. Each has its own permissions and compositor compatibility, so select the current session explicitly.
Keep the terminal focused until output completes and never have the pipeline send Enter. A safer alternative is standard output or clipboard review followed by a deliberate paste, especially for commands containing flags, paths or secrets.
What quality should you expect?
Recognition depends on the chosen Vosk model, microphone, language and environment. The project's limitations note lowercase output and possible cold-start delay, while larger models can improve accuracy at additional cost. Its Python configuration can post-process text, but rules do not understand every code identifier.
The Vibe Coding Pad provides three physical controls and a clip-on microphone, yet start with begin, cancel and reviewed paste. Keep the original transcript visible, test release failures and treat terminal text as a draft until every character has been checked. Save the exact audio and simulated-input backends alongside the model choice. When an update breaks the path, direct text to standard output first and test capture, recognition and injection as three separate stages before editing the hardware bind or changing permissions again.