A privacy page can be entirely truthful about not selling your data while staying quiet on a separate question: whether the recordings and transcripts that pass through the product are used to improve the model itself, for the benefit of every other user of that product. That's a legitimate business practice most vendors are upfront about somewhere, and it's also the specific question this page answers for seven Mac dictation tools, using each vendor's own current wording rather than a summary of what they probably mean. "Processed" and "trained on" get treated as synonyms constantly, and they aren't — an app can process your audio locally and never train on anything, or send it to a server and still never use it for training, or do both. The vendor's own stated policy is the only way to tell which.
Can't train on it — there's no channel for it to travel through
Coii VoiceInput
Model weights are fetched once, on first run, from a public model host. After that, the only network activity is a licence check when you enter a key and a clock check at launch — nothing during an actual dictation. No audio or transcript is ever sent anywhere, so there's no version of "trained on your voice" that could apply: training requires the data to leave the machine and reach whoever would do the training, and here it never does.
- 1Key downthe target is chosen here
- 2You speaklevel meter, over your work
- 3Key up
- 4Written to historybefore the engine is asked
- 5Engineon your Mac
- 6Typedabout half a second
VoiceInk
Checked on tryvoiceink.com today: transcription runs locally, and the app is open source, which is a stronger assurance than a written policy — the claim is, in principle, independently checkable by reading the code rather than only assertable. Its site does not make a specific statement about model training, which for a local-processing, open-source tool is a reasonable silence rather than a red flag. The comparison has the rest.
Opt-in, off by default — a stated, specific policy
Apple Dictation
Checked against Apple's own privacy documentation today: whether your Mac processes Dictation locally or on a server depends on your device and is shown in Keyboard Settings, but either way, "unless you opt in to Improve Siri and Dictation, your audio data is not stored by Apple," and for Dictation specifically, "the things you dictate are sent to and processed on the server, but will not be stored unless you opt in." That opt-in is off by default. If you do turn it on, Apple states it may retain transcripts for up to two years, tied to a rotating identifier not linked to your Apple Account, to improve Siri, Dictation and related features — a detailed, specific policy, and one where the default is genuinely off rather than a checkbox most people never find. The comparison has the rest.
Not stated — which is not the same as "yes"
MacWhisper
Checked on macwhisper.com today: the site describes local transcription and a local dictation mode, with no statement either way about using recordings or transcripts to train anything. For a tool whose core pitch is running on-device, that silence is a reasonable inference in the app's favour, but it's still a silence rather than a sentence you can quote — worth knowing the difference before assuming. The comparison has the rest.
superwhisper
Checked on superwhisper's Pro documentation today: local voice models are listed among what a Pro licence includes, and the free tier includes some cloud models — meaning the tier and model you pick determines whether audio leaves the Mac at all. Nothing on their site states a model-training policy either way. If this question matters to you specifically, confirming which model you're running is the practical next step, not assuming from the "local-first" framing alone. The comparison covers the rest of a much larger product.
BetterDictation
Checked on betterdictation.com today: core dictation runs on-device with no cloud round-trip, on Apple silicon only. An optional Pro add-on at $2/month states plainly that it "uses OpenAI" for post-processing, meaning transcribed text — not raw audio — leaves the Mac if and only if that feature is turned on. Their site makes no separate statement about model training for either path. The comparison sets out the rest.
Opt-out, on by default — the honest read of the wording
Wispr Flow
Checked on wisprflow.ai/pricing today: every tier, including Free, lists "Opt out of model training at any time" as a feature. The wording matters: an opt-out is only worth stating if the default is on, so read plainly, training is the default state until you change it yourself. Their site does not publish, on that page, how long opted-in data is retained or exactly what it's used to train. Pro is $15 a month, on Mac, Windows, iOS and Android, with a learned vocabulary that syncs across devices — a real advantage that this training question doesn't cancel out on its own. The comparison covers the rest of a considerably more capable product on most other axes.
The table
| Can data reach the vendor at all | Stated training default | How it's verifiable | |
|---|---|---|---|
| Coii VoiceInput | No — nothing sent | N/A | No channel exists |
| VoiceInk | No, local | Not stated | Vendor statement, open source |
| Apple Dictation | Depends on device/setting | Opt-in, off by default | Vendor statement, specific |
| MacWhisper (local) | No | Not stated | Vendor statement |
| superwhisper | Tier/model-dependent | Not stated | Vendor statement |
| BetterDictation | No (stated exception for optional add-on) | Not stated | Vendor statement |
| Wispr Flow | Yes, by design | Opt-out, on by default | Vendor statement |
Why "not stated" is worth reading as its own category
A vendor's silence on model training isn't evidence of anything specific — it could mean a policy exists and simply isn't on the marketing site, or it could mean the question genuinely doesn't apply to a tool that never sends data anywhere. What it isn't is a "no," and treating a silent page the same as an explicit, opt-in, off-by-default statement like Apple's overstates what you actually know. Three separate levels of confidence sit on this page: architecturally impossible, explicitly ruled out by policy, and simply unaddressed. They deserve different amounts of trust even when the vendors involved are all acting in good faith.
Who this actually matters most to
Anyone dictating client names, case details, unreleased product plans or anything else they wouldn't want quietly feeding a model other users draw from has a real reason to prefer the left side of this page's spectrum over the right. Everyone else — dictating grocery lists, calendar invites and Slack replies — is choosing between reasonable options either way, and a stated, honest opt-out like Wispr Flow's is not a mark against a product that's otherwise significantly more capable than the alternatives here.
What "training" actually changes, for people who aren't you
It's worth being specific about why this question is different from the general privacy question of whether your own data is safe. When a recording or a transcript is used to train a model, the thing that's changed as a result isn't just a record about you sitting on a server somewhere — it's the behaviour of the product every other user of that vendor's tool experiences afterward. That's not inherently a bad trade: a model that improves because real usage fed back into it is, in the aggregate, a reasonable way for a product to get better over time, and it's how most consumer AI products operate today. The reason it's worth separating from "is my data stored safely" is that the two questions have different answers even for a single vendor — data can be stored securely and still be used for training, or never stored at all and therefore never available to train anything, which is architecturally the stronger guarantee of the two.
A pricing page and a privacy policy are different documents
Most of what's quoted on this page comes from a pricing or feature page rather than a full privacy policy, because that's where vendors state the plain, one-line version of their training stance — "opt out any time," "processes locally." A full privacy policy usually says more, including retention windows, subprocessors, and the legal basis for processing, and it's worth reading directly if a specific compliance requirement is on the line rather than general curiosity. What a comparison page like this one can responsibly do is quote the plain-language claim a vendor puts in front of a buyer and link to where they said it, which is what's above; it isn't a substitute for reading the underlying policy yourself when the stakes are higher than routine writing.
What happens if a vendor changes hands or changes its mind
A stated policy, even an honest and currently-followed one, describes today's intent under today's ownership and today's regulatory environment. Any of those three can change, and a policy is generally the first document rewritten when one does — not necessarily in bad faith, but because a new owner, a new law, or a new feature that needs more data than the last one routinely triggers an update, and users are notified rather than consulted. An architecture where the data was simply never transmitted anywhere doesn't carry that exposure, because there's no data sitting anywhere for a future policy to reinterpret. That's the structural reason "can't" and "currently doesn't" are different strengths of guarantee, even when every vendor involved is acting in complete good faith today.
What to actually check before trusting any of this
None of the vendor statements quoted on this page are secondhand — each one links to the page it came from, and the honest next step for anyone whose decision genuinely rides on this is to read that page directly rather than trust a summary, this one included. Policies get updated, pricing pages get redesigned, and a claim that's accurate today is only guaranteed to be accurate today. For a use case where the training question is a compliance requirement rather than a preference, confirming the current wording on the vendor's own site, on the day the decision is made, is worth the five minutes it takes.
The difference a purchase model can make to the incentive
It's worth noticing, separately from any specific policy, that the business model behind a product shapes the incentive to train on user data in the first place. A subscription product priced to compete on capability has an ongoing reason to improve its model continuously, and training on real usage is one of the more effective ways to do that — which is a reasonable explanation for why the cloud, subscription tools on this page are the ones with an active training question to answer at all. A one-time purchase with no server component has less structural reason to want the data even before any policy is written, simply because there's no recurring product improvement loop the purchase is funding. That's not a claim that subscription pricing implies bad faith — Wispr Flow's stated, working opt-out is a real commitment — it's an observation about why this specific question comes up for some categories of product far more than others.
One more distinction worth naming
It's tempting to treat "doesn't train on my voice" and "doesn't listen to my voice at all" as the same claim, and for most of the tools on this page they amount to close to the same thing in practice. But they're not identical: a vendor could, in principle, process audio to produce a transcript, discard the audio immediately, and still never touch training — in which case the training answer is "no" even though a server briefly saw the audio. None of the vendors above are alleged to have a training policy that contradicts their processing claims; the point is only that the two questions are worth keeping separate when reading any vendor's wording closely, rather than assuming one implies the other.
Picking one
If the requirement is that training simply cannot happen, because there's nowhere for the data to go, the trial for Coii VoiceInput is thirty days of exactly that, with no setting to double-check. If you want the most capable dictation product on the market and are comfortable with an explicit opt-out you have to exercise yourself, Wispr Flow states its policy plainly and lets you turn training off. If open source matters more to you than a written promise, VoiceInk is the one entry here where the claim is, in principle, independently checkable.