The average typist manages around 40 words a minute. Ordinary conversational speech runs closer to 120 to 150. On paper that is a three-to-one advantage, and it is the whole reason speech-to-text on a Mac is worth looking at for anyone who spends a working day producing sentences — email, reports, messages, notes. In practice almost nobody gets the full gap, because the speed only counts once the words land correctly and stay landed, and that is where most tools quietly give it back.
What "speech to text" actually covers
The phrase gets used for two genuinely different jobs, and it is worth knowing which one you actually need before comparing prices. One is live dictation: speech turned into text as you say it, straight into whatever you are already typing into. The other is transcribing a file after the fact — a recorded meeting, an interview, a voice memo — into a document you did not have to type. MacWhisper does both: system-wide dictation on its free tier, plus file transcription for recordings, which makes it one app covering two different jobs. Coii VoiceInput only does the first — live dictation into whatever app you are using — and does not open or transcribe an existing recording at all.
Where the time actually gets lost
Two failure modes eat the speed advantage before it reaches you. The first is the tool getting a word wrong and forcing a correction — a proper noun, a piece of jargon, a name it has never heard — which is slower to fix by hand than it would have been to type in the first place. The second is rarer but worse: the transcription simply does not arrive where it was supposed to, because the app that was meant to receive it lost focus, or the process that turns speech into text stalled, and now you are retyping the whole sentence from memory instead of correcting one word.
Apple's own Dictation, free and built into every Mac, is a perfectly good answer to the first problem for plenty of people and says nothing about the second — the full comparison goes through exactly what it does and does not handle. What follows is aimed at both problems together.
What a speech-to-text tool has to get right
The transcript exists before the engine is even asked to produce text
The moment you finish speaking, Coii VoiceInput writes the words to a local history — before the transcription engine is asked anything. That inverts the usual failure order: normally the words only exist once the engine succeeds and the paste lands; here they exist first, and the engine and the paste are just how they also end up in your document.
- 14:025.8sThe Q3 numbers came in ahead of forecast across every region except EMEA.Pages
- 14:153.4sFollowing up — can we push the review to Thursday morning instead?Mail
- 14:318.1sDraft summary for the board: revenue up, churn flat, hiring paused through year end.Notes
The destination is decided before you start talking, not after
Whichever field is focused when you press and hold ⌥ Space is where the text goes, through the accessibility API where the app supports it and a synthetic paste where it does not, with the clipboard restored either way. There is no separate "insert" step to get wrong — the words appear where you were already working.
Nothing about accuracy depends on a network round-trip
The engine that turns your speech into text runs on the Mac itself, so there is no request to a server that can be slow, queued, or unreachable on a bad connection. Wispr Flow, one of the largest names in this category, transcribes on its own servers and needs a live connection every time; the tradeoff either way is speed and reliability against a larger, continuously updated model on someone else's machine.
It handles a whole working day of different writing, not one document
Reports, email replies, quick notes, Slack messages — the same held key works across all of them because it is a system-wide shortcut rather than a feature bolted onto one app. A longer list of apps built for long-form writing is worth reading separately if most of your day is one continuous document rather than many short ones — the apps that win there are not always the same ones that win for quick replies.
Punctuation stays a habit you build, not something spoken
Saying "comma" or "period" out loud tends to type the word rather than the mark, because nothing here distinguishes "the word comma" from "insert a comma" without a command grammar built specifically for that — and this app does not have one. The workaround almost everyone settles into after a short adjustment period is dictating the words alone and adding punctuation by hand afterward, which sounds like it defeats the speed advantage and in practice costs a few seconds per sentence rather than the minutes typing the whole thing would have taken.
What it does not do: learn your vocabulary
Said plainly, because it is the honest gap: there is no dictionary to teach it a project name or an unusual surname, so the same word comes out the same way every time, right or wrong. Wispr Flow and a handful of others build exactly that feature; the developer-focused round-up covers who has it and what it costs.
What accuracy actually depends on
Two things determine whether a spoken sentence comes out right, and neither is really about which app you bought. The first is the audio itself — a quiet room and a decent built-in microphone gets a noticeably cleaner result than a noisy café through a laptop mic held at arm's length, for any transcription engine, local or cloud. The second is vocabulary: plain conversational English is the case every engine is tuned hardest for, and accuracy drops in a fairly predictable order as a sentence gets further from that — proper nouns first, then acronyms and internal jargon, then anything genuinely unusual like a name from a language the engine was not trained heavily on.
None of that is specific to this app; it is closer to a property of speech recognition generally, and it is worth testing on your own voice and your own vocabulary before trusting any accuracy claim, including implicit ones made by omission on a comparison page like this one.
A working day, spoken instead of typed
The clearest way to see whether the math in this page's opening actually holds is to walk through an ordinary day rather than a single sentence. A morning email answered in the time it takes to say it rather than type it. A meeting follow-up dictated the moment the call ends, while the details are still fresh, instead of typed twenty minutes later from a half-remembered summary. A quick note to yourself about something to fix before end of day, said in passing rather than opened as a whole separate task. None of those are dramatic on their own — a few seconds here, a minute there — and the case for speech-to-text is really that those small moments add up across a week in a way a single demo sentence cannot show.
What changes about your writing process, and what does not
Speech-to-text changes where words come from, not what happens to them afterward. A sentence dictated into an email still gets read before it is sent; a paragraph dictated into a report still gets checked against the numbers it describes. Nothing about speaking instead of typing removes the editing step most people already do as a matter of habit — if anything, a first pass produced quickly leaves more time for that step rather than less, because it did not consume the time editing usually has to compete with.
The adjustment that actually takes getting used to is smaller and more mechanical: composing a sentence out loud, in order, without the ability to jump back and rework the middle of it the way a cursor lets you. Most people find that adjustment settles within the first week of regular use — long enough to stop noticing the difference, and to start noticing instead which parts of a normal day got faster.
What it costs and what it needs
$19 once, three devices, macOS 13 Ventura or later on Apple Silicon or Intel — no subscription, no account, updates included. The thirty-day trial runs the full app, so the speed-versus-corrections tradeoff above is something you can actually measure on your own writing before paying anything.
Who this actually helps
If most of what you write in a day is short and low-stakes — quick replies, brief notes — the time you would save is small and probably not worth learning a new habit for. The general case for a dictation app on this site covers that reader more directly.
If your day is full of longer prose — reports, summaries, correspondence, the kind of writing where forty words a minute genuinely feels slow — the gap between speaking and typing is large enough to be worth thirty days of testing it for free against your actual writing, not a demo sentence. And if speed is only part of the reason you are here — if where the audio goes matters as much as how fast it becomes text — that is worth weighing before price does, since it is the one property a faster cloud engine cannot offer back no matter how good its accuracy gets.
The most reliable way to answer the question this page opened with — does speaking actually save you time, once corrections and lost sentences are counted — is to measure it yourself for a week rather than take either a marketing claim or a skeptic's anecdote on faith. Pick one recurring task, a daily status update, a set of email replies, a section of a longer document, and do it by voice for five working days running. Most people who try that honestly either notice the time back within the first two or three days, or notice clearly that the friction outweighs it — and either outcome is more useful than a percentage on a page like this one, because it is measured against your own sentences instead of somebody else's benchmark. A week is short enough to run alongside a normal workload and long enough to get past the first day's unfamiliarity, which is usually where most of the friction actually lives before a new habit settles in. By the end of it, the decision tends to make itself, one way or the other, without needing another comparison page to weigh in on your specific working day, your specific vocabulary, or the specific mix of short replies and longer documents that actually makes up your week.
Attaching the revised proposal — happy to walk through the pricing changes on a call this week.