§ Feature
Voice input and dictation for macOS
Hold a key, speak, let go. Your words land as text wherever the cursor is — or, held over a selection, rewrite what is already there.
Hold-to-talk, not always-listening
Press and hold the trigger — Fn, Right ⌥, or a Persona’s own hotkey — and KeyMic starts capturing. Release, and it stops recording but keeps listening for up to 2 seconds in case the recognizer is still catching up on your last words, then delivers whatever it has. There is no wake word and nothing runs while the key is up. A 6-minute cap ends any session that outlives a forgotten key release, and pressing any other key mid-capture cancels it outright rather than transcribing a half-formed thought.
While you talk, a small overlay shows the live partial transcript so you can see what KeyMic is hearing before it lands.
Four engines, one picker
Voice input is not tied to one recognizer. Pick a model in Settings and every session uses it until you change it again:
- Apple (system) — the built-in
SFSpeechRecognizer, covering every language macOS itself supports, no download required. - SenseVoice Small — a fully local, on-device model (≈ 198 MB) covering Mandarin, Cantonese, English, Japanese and Korean.
- Fun-ASR-Nano — a local model (≈ 1 GB) for Mandarin, English and Japanese.
- Fun-ASR-MLT-Nano — the same local runtime tuned for 31 languages, for the widest single-model coverage without leaving the device.
Local models require macOS 15 or later; on macOS 26 the Apple path itself can upgrade to Apple’s newer SpeechAnalyzer engine automatically once your language and its on-device assets are ready — you keep using the same “Apple (system)” option, KeyMic just picks the better backend underneath it.
Reshape text with a Persona
Select text, hold a Persona’s hotkey instead of the plain voice trigger, and say what you want done — “make it formal,” “translate to English,” “uppercase this.” KeyMic sends the selection and your instruction together and replaces the selection with the result, rather than just inserting a transcript.
Where the limits are
One engine is active per session — KeyMic does not blend models or auto-detect language across an utterance; you choose a model, and that model’s language list is what it understands until you switch it in Settings. Delivery works by placing the result on the pasteboard and pasting it, then restoring whatever was on your clipboard before — quick and format-preserving, but it does mean a paste, not raw keystrokes, is what lands in the target app.
Permissions
Voice input needs Microphone access, requested the first time you use it. Local models also need one-time downloads of their weights before they can run offline; the Apple option needs nothing beyond system Speech Recognition permission. Security & Architecture covers what stays on-device for each engine.
Also in KeyMic: hotkey-driven actions, or see everything on one page.