Skip to On this page All posts
dictationmacoslive-transcriptionwhisperon-device

Live Transcription That Keeps Up With You

EnviousWispr now transcribes while you speak, so long dictations finish almost instantly when you stop. Here is how it works, and why Auto-detect waits.

The slowest part of dictation is not the talking. It is the moment after you stop, when you are waiting for your words to show up.

Until now, EnviousWispr’s All Languages engine handled that moment the traditional way: it collected your audio while you spoke, and when you released the key, it transcribed the whole thing in one pass. That works, and it is accurate. But the wait grows with the length of the dictation. Talk for two minutes and you wait for two minutes of audio to be processed. Talk through a whole idea, the way an hour-long session invites you to, and the pause at the end starts to pull you out of your flow.

In the latest update, the All Languages engine transcribes while you speak. When you stop, there is almost nothing left to do, so your text lands almost immediately, whether you spoke for ten seconds or ten minutes.

How it works

While you are talking, the engine quietly transcribes the audio it has so far, over and over, every couple of seconds. Each new pass hears a little more than the last one.

Here is the part I find genuinely elegant. A word is only committed once two consecutive passes agree on it. If the engine hears “let’s meet at the” on one pass and the next pass hears the same thing, those words are locked in. If the passes disagree, the words stay tentative, and the engine simply waits for more audio to make up its mind. As full sentences get confirmed, they are set aside as finished, so each new pass only works on the recent, still-open part of your speech and stays fast no matter how long you talk.

This approach comes from a research group at Charles University that studies live speech translation, where the same problem shows up in a harder form. Their method, called whisper_streaming, is the design we adapted. We benchmarked it against other candidates on real dictation before picking it; it won on both speed and text quality.

By the time you release the key, nearly everything you said is already confirmed. The engine finishes the last unconfirmed words, and that is it. The end-of-dictation wait stops scaling with how long you spoke.

And if anything goes wrong mid-recording, the engine keeps your full audio on the side and falls back to transcribing it in one pass, the traditional way. You always get your text. The fast path is never allowed to put your words at risk.

What Auto-detect language changes

There is one setting that turns this off on purpose: Auto-detect language.

Live transcription has to commit to a language on its very first pass, a second or two into your recording. That is almost no audio to judge from. If it guessed wrong there, every confirmed word after that would be in the wrong language, and confirmed words cannot be taken back.

So with Auto-detect on, EnviousWispr does the careful thing instead: it waits until you finish, then judges the language from your opening audio, seconds of it, rather than the split second live mode would have to decide from. You get the traditional end-of-recording wait, and in exchange you get the right language essentially every time. The app notes this right under the toggle, so it is never a mystery.

If you dictate in one language, the fix is simple: pick it. Settings, Transcription, Language. Locked to a specific language, the All Languages engine streams live.

Here is the full picture:

Live transcriptionLanguage settingWhat you get
OnA specific languageTranscribes while you speak, near-instant finish
OnAuto-detectTranscribes at the end, best language accuracy
OffEitherTranscribes at the end

Why you might still turn it off

The toggle exists for a reason. On very long recordings, a single pass over the finished audio can produce slightly cleaner text, because the engine gets to hear everything with full context before writing anything down. If you regularly dictate long-form and care about every comma, try both and see which reads better for you. For everyday dictation, live transcription produces the same text and gives you the time back.

Still on your Mac, in both modes

None of this changes where the work happens. Live or not, transcription runs entirely on your device, with no upload, no server, and no account. The audio from the microphone goes to the model running on your Mac’s own chip and nowhere else. Live transcription changes when the work happens, not where.

If you have ever wondered why we care so much about squeezing speed out of local models, from moving Whisper onto the GPU to this, it is because on-device is only a real alternative to the cloud if it also feels instant. This is one more piece of that.

Update to the latest version, or download EnviousWispr free if you have not tried it. Hold the key, talk as long as you like, and watch how little you wait.

Try EnviousWispr free. On-device dictation for Mac, no account required.

Download Free