Verbatim dictation means the text on screen contains the words that came out of your mouth: the hedges, the false starts, the "um" if you said one, the sentence that trails off. It is what a journalist needs from a quote, what a researcher needs from an interview note, what anyone dictating a prompt needs, and what most paid Mac dictation apps will not give you without a settings change — because tidying your speech is what they are sold on.
This page covers how to get it: checking what your current app does, switching off what can be switched off, and what to do about the parts that cannot.
Step 1 — Find out what you currently have
Before changing any setting, establish the baseline. Open a plain text editor and dictate these three things exactly.
The hedged sentence with a correction.
I think we should probably ship on Thursday — actually, Friday — assuming the tests pass.
Read the output. Did "I think" and "probably" survive? Is "Thursday" still there, or did the app quietly resolve the correction for you?
The one with fillers.
Um, so the thing is, uh, we already tried that and it didn't work.
Count the ums on screen.
The run-on. Speak sixty words without pausing and without saying any punctuation. If neat sentences come back with commas and full stops, the app made structural decisions about where your thoughts ended.
Whatever you learn here applies to everything you have dictated so far, which is the uncomfortable part. Does a dictation app change what you said covers the vocabulary to look for in a feature list, so you can confirm the finding against what the developer says it does.
Step 2 — Turn off what the app lets you turn off
The settings that matter have fairly consistent names across the category. Look for and disable:
- Filler word removal — sometimes "remove ums", sometimes bundled into a cleanup toggle with no separate switch
- Auto cleanup, polish, enhance
- Grammar correction or grammar improvement
- Formatting, auto-format, formats as you speak
- Tone, style, register, or a mode picker offering formal / casual
- Modes aimed at email or messages, which usually reformat for that medium
Two things to know before you go looking. First, in several apps the editing pass is the product rather than a feature, so there is no switch — the only way to get verbatim output is a different app. Second, some tools keep a "raw" or "original" view in their history even when cleanup is on, and if yours does, that is the fastest route to a verbatim copy of something you already dictated.
Step 3 — Set up an app that does it by default
Two options on a Mac need no configuration at all.
Apple's built-in dictation transcribes what it hears without an editing pass. It is free and already installed, and for someone who only needs verbatim capture occasionally it may be the whole answer — Apple Dictation against a paid app is honest about where it stops being enough, which is mostly about where the text lands and what happens on a long dictation.
Coii VoiceInput ships with filler removal off. Its post-processing is a fixed chain — repeat collapsing, your own dictionary, optional filler removal, punctuation and spacing — and no model is asked to improve the sentence. The chain has no stage that can turn one phrase into a different phrase, which is the property that makes verbatim output reliable rather than a setting that might drift.
To confirm the setting on a fresh install: open Settings, find Post-processing, and check that filler removal is unticked. It is unticked by default and is intended to stay that way.
Um, the second act is, I think, still too slow — or maybe it's just the opening.
Step 4 — Keep the original available
Even a verbatim setup runs some filters, and a dictionary entry can fire where you did not want it. The safeguard is whether the app keeps what it originally heard.
Coii VoiceInput stores both the raw transcript and the processed one on every history row. If a filter did something you disagree with, open history and take the raw version; the original is never overwritten by a filter's opinion of it. That is what makes it safe to have filters at all, and it is worth checking whether the app you use does the same — if it stores only the output, a bad edit is unrecoverable.
The transcript is written to history before the engine is asked anything else, so a slow model or a paste that never lands cannot cost you the words either.
Step 5 — Punctuation is a separate decision
Verbatim is about phrasing, not about whether a full stop exists. Mechanical punctuation does not change which words you used or what they mean, so leaving it on is compatible with verbatim capture, and most people want it.
Where you do want full control, speak the punctuation — "comma", "full stop", "new paragraph" — and turn the automatic pass off. Where the text ends up matters as much as how it is punctuated, and speech-to-text on a Mac covers the mechanics of getting it into the field you were already typing in.
What verbatim does not mean
Two things get filed under verbatim that are not the same thing, and confusing them leads people to disable settings they actually wanted.
It does not mean no dictionary. If you have told the app that your company's product is spelled a particular way, an entry firing on that word is not the app overriding you — it is the app doing what you typed into it. Your own substitutions are the one category of change you authored, and they are worth auditing when something looks wrong: a short entry can match inside a word you did not intend, and the fix is a more specific entry rather than turning the dictionary off.
It does not mean no formatting conventions. A capital letter at the start of a sentence and a space between a Latin word and the Chinese characters around it are typographic conventions, not edits. They do not change which words you used, and nobody reading the output could reconstruct a different meaning from them. Keeping them on costs nothing that verbatim is trying to protect.
What verbatim does mean is narrow and worth stating precisely: the sequence of words in the output is the sequence of words you spoke. Nothing was removed for being untidy, nothing was added for smoothness, and nothing was swapped for a better-sounding equivalent. Everything else — punctuation, casing, spacing, the entries you typed yourself — is mechanics.
The other thing that is not verbatim: the engine repeating itself
Editing passes are the obvious enemy of verbatim capture. There is a second and less discussed one, which is the engine producing words nobody said.
Speech recognition decoders occasionally loop. A phrase comes out three times, or a run of text repeats to the end of the utterance, usually triggered by silence, background noise or a microphone that dropped out mid-sentence. This is not an editing decision — it is the recogniser failing — but from your side it has the same effect: the transcript is not what you said.
Coii VoiceInput collapses those repeats as the first stage of the chain, before anything else reads the words, on the basis that a decoder loop is unambiguously wrong in a way a hedge is not. That is the one place the app deletes text you did not ask it to delete, and it earns it by being the one case with no legitimate reading.
If you are chasing down a verbatim problem, work out which of the two you have. Text that is missing was probably edited; text that is duplicated was probably looped. They have different fixes, and a cleaner input — a headset rather than the built-in mic, and a room without a fan in it — helps the second one and does nothing for the first.
Verbatim in two languages at once
Anyone who mixes languages in a single sentence on purpose has a specific version of this problem: a cleanup pass frequently decides one of the languages was a mistake. A technical term in English inside a Chinese sentence, a loanword left in its original form, a name that only exists in one script — all of these are exactly what a grammar corrector is built to normalise away.
Verbatim output is the only setting under which mixed-language dictation survives intact. The mechanical filters that stay on are conventions rather than corrections: a space between scripts where the convention calls for one, full-width punctuation in Chinese text. Neither changes which words you used. Multilingual dictation covers which tools handle this at all, which is a shorter list than the hundred-language marketing suggests.
A workflow that gets both
The usual objection to verbatim dictation is that you end up editing more, and that is true. The way out is not to pick one mode for everything but to keep the two artefacts separate.
Dictate verbatim. Let the transcript be the record — messy, hedged, exactly what you said. Then edit it yourself into the version you send, or hand it to whatever tool you already use for drafting, with the original still sitting in history where you can check it.
The value of that split is that the two stages are legible. When the polished version says something you did not mean, you can see where it came from. When the app does the two silently and in one step, the only thing you have is the result, and no way to tell whether the meaning moved during transcription or during tidying.
For a long piece this is also less work than it sounds. A morning of dictation is one editing pass either way; the difference is whether you are editing your own sentences or an app's interpretation of them.
Who this actually matters for
Journalists. A quote that has been grammar-corrected is not a quote, and a verbatim note is the difference between something you can publish and something you have to go back and verify. Dictation for journalists covers the wider field-notes workflow.
Writers. If your prose has a voice, an editing pass reliably normalises it toward well-formed and less like you. The rough sentence is the one worth keeping, because you can edit it yourself later. Dictation for writers goes further into this.
Anyone dictating prompts. Your exact phrasing is the input. A tidying pass before the text reaches the assistant changes the instruction, and you will be judging the result two steps downstream without knowing the input moved. Dictating prompts to an AI assistant is the page on that.
Anyone whose hedges are load-bearing. Clinical notes, legal memos, code review, status updates. "This should work" and "this will work" are different claims, and only one of them is yours.
A note on microphones
Verbatim capture is only as good as the audio, and the single biggest source of words you did not say is a recogniser guessing at a signal it could not hear properly. A built-in laptop microphone in a room with a fan, an air conditioner, or someone else's conversation in it will produce plausible invented text rather than silence, because a decoder asked to find words in noise finds words.
A headset microphone close to your mouth removes most of that, and it is the cheapest accuracy improvement available in this category — cheaper than any app, including this one. If you dictate for hours a day and have never changed the input device, that is the first thing to try, before concluding that any particular tool hears you badly.
What it costs you
Verbatim output is messier and you will edit more. That is the trade and it is a real one: if most of what you dictate is email to people you do not know well, an app that produces a clean formatted paragraph is saving you a pass on every message. The argument for the other side, including who should buy which, is in dictation that doesn't rewrite your words.
Coii VoiceInput is $19 once, macOS 13 or later, with a 30-day trial and no account — the trial is long enough to find out whether you miss the tidying, which is the only test that settles it.