Key findings
- Otter.ai uploads your audio and keeps it. Deleted recordings sit in trash for 30 days, and the policy reserves a right to retain material past deletion for legitimate business purposes.
- Otter's AI training setting is opt-out. New accounts contribute until the user finds the toggle.
- Granola discards audio but still uploads it first. Recordings go to third-party processors for transcription, then transcripts are stored in AWS.
- Fathom keeps the most. Its record can include audio, video and clips, which is a feature for sales teams and a liability for everyone else.
- Only on-device transcription avoids the upload entirely, and it costs you cloud-quality summarisation. Nobody currently gives you both.
We build a transcription app, so we read a lot of competitor privacy policies. The pattern is consistent: the marketing page says "enterprise-grade security" and "your data is yours," and then the policy says the recording goes to a cloud bucket, gets passed to two or three subprocessors, and may outlive the delete button.
None of that is unusual or sinister. It is just how cloud transcription works. The problem is that the people most likely to be recording sensitive conversations, which is to say therapists, journalists, lawyers, doctors, HR staff and anyone doing user research under an NDA, are the least likely to read fourteen pages of policy before hitting record.
So this is that reading, done for you, on the only question that changes your risk: where does the audio go.
The three questions that matter
"Private" gets used for three completely different architectures, and conflating them is how people end up surprised.
Does the raw audio leave the device? This is the big one. Audio is biometric-adjacent data, it identifies speakers, and it captures everyone in the room, including people who never agreed to your software choice. A tool that never uploads audio has a categorically smaller blast radius than one that does, regardless of how good the second one's encryption is.
Who else touches it? Most transcription products are not doing their own speech recognition. They pass your audio to a specialist provider, and the summary to a model vendor. Each hop is another company, another policy, another jurisdiction, another breach surface.
What survives deletion? Two separate things: retained copies, and model weights. If your recording was used for training before you deleted it, the deletion does not reach into the model.
Every claim below comes from the vendor's own published policy or documentation as of August 2026, linked inline. Policies change often. Granola's current policy, for instance, took effect on 24 July 2026. If you are making a compliance decision rather than a personal one, read the current version yourself rather than trusting any comparison article, this one included.
The comparison
| Tool | Raw audio uploaded | Audio kept after transcription | Third parties in the path | Training default |
|---|---|---|---|---|
| Otter.ai | Yes | Yes, 30-day trash window, plus retention rights past deletion | AWS and data labeling providers | Opt-out |
| Granola | Yes | No, discarded once the transcript exists | Deepgram or AssemblyAI, OpenAI and Anthropic, AWS | See current policy |
| Fathom | Yes | Yes, records can include audio, video and clips | Cloud and model vendors | See current policy |
| Voxxli | No | Stays on the phone | One AI provider, transcript text only, see below | Not used for training |
Two honest caveats about that table. First, it covers four tools, not the whole market, because we would rather publish four rows we have verified than nine we have half-checked. Second, the last row is our own product, and it has a compromise in it that we describe in full below rather than burying.
Otter.ai
Otter is the default in this category and it is genuinely good at the job it does. It is also the least private option here by some distance.
Audio uploads to Otter's cloud immediately on recording, and is shared with third-party service providers including AWS. That much is ordinary. The parts worth knowing:
- Training is opt-out. New accounts contribute to model training until the user goes and finds the setting. Otter also uses data labeling providers, meaning human annotators, who build training and evaluation sets from shared data.
- Deletion is soft for 30 days. Deleted items go to trash and are permanently removed after that window.
- The policy reserves retention past deletion where Otter determines there is a legitimate business purpose, and it does not commit to specific retention periods elsewhere.
- Training data outlives the source. Information already absorbed into a model can persist there after you delete the recording it came from. This is true of every vendor that trains on customer data, and it is the reason the opt-out default matters more than the delete button.
If you are recording your own standups, none of this should worry you. If you are recording a patient, a source, or a candidate, the opt-out training default alone is probably disqualifying, and you want to check that setting before your first session rather than after.
Read it yourself: Otter.ai privacy policy.
Granola
Granola is the most interesting one here, because they clearly thought about this and made a deliberate tradeoff in the other direction.
Audio is captured locally on macOS and Windows, then sent to cloud subprocessors for transcription, reported as Deepgram and AssemblyAI, with OpenAI and Anthropic handling summarisation. Transcripts are stored in AWS. Crucially, recordings are captured for transcription only and are not retained once the transcript exists. The audio is discarded. Transcripts and notes persist until you delete them or your organisation's retention policy removes them.
Granola has also said publicly that they tried on-device transcription and moved it to the cloud for quality reasons. We believe them, and it is worth sitting with that, because it is the actual tradeoff in this entire category. Cloud models are better. On-device models are private. Anyone claiming to give you both at full strength is selling something.
Granola's position is defensible: upload the audio, use the best available model, throw the audio away immediately. If your concern is long-term storage of your voice, that solves most of it. If your concern is that the audio touched four companies' infrastructure at all, it does not.
Read it yourself: Granola privacy policy, effective 24 July 2026.
Fathom
Fathom is built for sales teams, and it is designed around a different assumption: the recording is the deliverable. Where Granola preserves the transcript and discards the audio stream, Fathom preserves a record that can include audio, video and clips, because the point is to send your colleague the ninety seconds where the prospect said the interesting thing.
That is a real feature, and for a revenue team it is the correct design. It also means Fathom is the wrong tool for a conversation you would not want replayed. Choose it knowing which of those two situations you are in.
Record without the upload
Voxxli records and transcribes on your iPhone using Apple's on-device speech engine. No account, nothing to sign up for, and the audio file never leaves your phone. Free on the App Store.
Voxxli, and where it compromises
We built Voxxli, so treat this section with the scepticism it deserves. Here is the architecture, including the part that is not fully local.
What stays on your phone: recording, transcription, storage, search and export. Transcription runs on Apple's on-device speech framework. There is no upload step, and no account, so there is nothing on our side to breach. You can verify this in about five seconds by turning on airplane mode and recording something. The transcript still appears.
What leaves your phone: if you ask for a summary, flashcards, speaker labels or chat, the transcript text is sent to a third-party AI provider for processing. Not the audio. The audio file stays on the device in every case. The provider is named in our privacy policy, along with what they may and may not do with it.
That is a real compromise and we would rather state it plainly than let the marketing imply otherwise. Text is less sensitive than audio: it does not carry voice, tone, or the identity of everyone else in the room. But it is not nothing, and if a conversation is sensitive enough that even the text should not leave, use Voxxli for the recording and transcript and skip the AI features. The transcript is fully usable on its own.
How to check any tool in ten minutes
This article will go stale. The method will not. When you are evaluating any transcription product:
- Run the airplane mode test. Turn off all networking and record thirty seconds. If a transcript appears, transcription is genuinely local. If it spins or errors, it is not, whatever the landing page says.
- Search the privacy policy for "subprocessor" or "service provider." Vendors are usually required to name them. Count the companies. That is your real exposure surface, not the one company whose logo is on the app.
- Search for "retain" and "delete" and read every hit. You are looking for two things: a specific retention period, and any clause letting them keep data past deletion. Vague language here is itself the answer.
- Find the training setting before you record anything. Note whether it defaults on or off. A vendor that defaults it on has told you their priorities.
- Check whether deletion is soft or hard, and how long the window is.
Five checks, ten minutes, and it works on tools that did not exist when this was published.
Questions people actually ask
Does Otter.ai store your recordings?
Yes. Audio uploads to Otter's cloud infrastructure and is stored there, with AWS named among its service providers. Deleting moves an item to trash, where it is permanently removed after 30 days. The policy also reserves a right to retain material beyond deletion where Otter determines there is a legitimate business purpose.
Is any AI meeting transcription fully on-device?
Transcription can be. Summarisation generally is not. Apple's speech framework transcribes locally on iPhone with no network at all, which is what Voxxli uses. But large language model summaries, chat and speaker labelling almost always require sending text to a hosted model. A tool that keeps audio local and sends only transcript text is meaningfully more private than one uploading raw audio, and it is still not fully local. Be suspicious of anyone claiming otherwise.
Does Granola process audio on your device?
It captures audio locally on macOS and Windows, then sends it to cloud subprocessors for transcription, reported as Deepgram and AssemblyAI, with OpenAI and Anthropic for summaries. Granola states recordings are not retained once the transcript is created, and transcripts are stored in AWS. Granola has said it tried on-device transcription and moved to the cloud for quality.
What is the most private way to record an in-person meeting?
Record locally with a tool that transcribes on-device, then decide separately whether to send the text anywhere. The practical test is the airplane mode test above. If transcription works with networking off, the audio is being processed on your hardware.
Is Otter.ai opted into AI training by default?
Its training setting is opt-out, so new accounts contribute until the user changes it. Otter also uses data labeling providers who build training and evaluation sets from shared data. Anything already absorbed into a trained model can persist there after you delete the original recording.