A Push-to-Talk Setup for whisper.cpp

WhisperKey is a clear polycarbonate handheld voice keypad with mic, tick, cross and enter controls. It can provide the down and up events for a local whisper.cpp pipeline, but the transcription project does not supply a finished desktop dictation workflow by itself.

What does a local whisper.cpp pipeline need?

It needs microphone capture, a stop condition, a compatible model, inference, output cleanup and a reviewed way to place text in the destination. whisper.cpp includes command-line examples and a microphone stream example, but that stream tool is described as naive real-time inference rather than a universal dictation app.

The official stream example depends on SDL2 for capture and can run continuously or use a basic voice-activity mode. For deliberate push-to-talk, a wrapper can instead record one audio file and invoke the command-line transcriber after release. Pick one architecture before connecting hardware.

How can a held key control record and transcribe?

Use an operating-system hotkey layer that exposes separate key-down and key-up callbacks. Down starts one recorder process and stores its process identifier. Up stops that exact process, waits for the audio file to close and starts transcription. Guard against a second down event while a job is active.

Write states such as idle, recording, transcribing and ready to a small log or visible indicator. If the release event is lost, provide a separate cancel key that stops capture without inserting text.

How should text reach the focused window?

The safer design writes the transcript to a temporary file or clipboard and shows it for review. A second deliberate action can paste it. Simulated typing depends on accessibility permissions, keyboard layout and whatever window has focus after inference completes.

Never append Enter automatically. The focused window can change while the model runs, especially on slower hardware. Store no sensitive transcript longer than necessary and restrict temporary-file permissions.

What are the tradeoffs versus a packaged app?

Local inference gives direct control over models, audio files and network use, but installation, microphone selection, latency and text insertion are your responsibility. Speed and accuracy vary with model and hardware. The stream example's voice detector also needs tuning for the actual environment.

The WhisperKey supplies four labelled controls, not the software pipeline. Begin with mic for capture and cross for cancel, keep tick as a separate paste action and leave enter unbound until repeated tests prove the destination and state handling. Pin the tested model and executable version in your notes. A wrapper built around one command-line output format can fail silently when flags or parsed text change after an upgrade.

Back to blog