Research
Which AI notetakers train on your meeting data?

Dylan de Heer
There is one question that decides whether a notetaker is allowed anywhere near a privileged, clinical or NDA-covered conversation, and it is not about accuracy or price.
Does the vendor use your meetings to train its models?
The answers are more varied than you would expect, and they are rarely on the pricing page. Some vendors say no in a single sentence. Some say yes, on de-identified content, with a setting you can turn off. Some scope the answer so tightly that you have to read three documents to work out what it covers. And a few remove the question entirely by never receiving the recording in the first place.
Here is where each one stands, taken from the vendors' own pages, with the wording that matters.
The short version
Tool | Trains on your meetings? | How you turn it off | Where it says so |
|---|---|---|---|
Weeve | No. Nothing is uploaded, so there is nothing to train on | Not applicable | On-device by architecture |
Meetily | No. Local by default | Not applicable | Open source, MIT licensed |
MacWhisper | No by default, files stay on the Mac | Do not enable cloud transcription | Local processing by default |
Jamie | No, and states this covers its model providers too | Not applicable | Security page |
tl;dv | No | Not applicable | Security page |
Fireflies | No, with zero-day retention at its vendors | Not applicable | Security page |
Zoom | No, for its own and third-party models | Not applicable | Terms of service |
Fathom | Yes, de-identified, unless you opt out | Account settings | Privacy policy |
Granola | Yes, de-identified, unless you opt out. Enterprise is off by default | Account settings, on every plan | Privacy policy |
Otter | Yes, de-identified | Documented opt-out on Enterprise | Privacy policy |
Three groups, and the difference between them is not degree, it is kind.
The tools that say no are making a promise you have to trust, backed by a policy and in some cases a certification. That is a normal commercial arrangement and there is nothing wrong with it.
The tools that say yes with an opt-out are asking you to configure your way to the outcome you want. The setting is real and it works. The question is whether you are willing to rely on every colleague finding it.
The tools that never receive the recording are not making a promise at all. There is no policy to trust because there is no copy to govern.
The ones that train on your meetings by default
Fathom
Fathom's privacy policy states that it "may use and create de-identified data generated from Meeting Content Information to improve our Services by training, improving, and customizing our in-house artificial intelligence models", and that you can opt out in your account settings. Team Edition adds an organisation-wide opt-out.
It also draws a distinction worth reading carefully: separately, it states that it does not authorise third parties such as OpenAI, Anthropic and Google to use your information to train their models. Those are two different commitments. The first is about Fathom's own models and is on by default. The second is about everyone else's and is a flat no.
Fathom's help centre also states that all Fathom data is stored in the United States.
Granola
Granola's privacy policy states that it only uses de-identified data to train AI models, and that you can opt out in your account settings. The opt-out is available on every plan, including the free one, which is better than most vendors offer.
One exception worth knowing: Enterprise workspaces are opted out of training by default. So the answer depends on which Granola you are on.
Granola is also facing a proposed class action, Chamberlain v. Granola, filed in California federal court in July 2026, which alleges that it captures participants without their knowledge and that training is on by default on the Free and Business plans. As with the Otter case below, those are allegations in a filed complaint rather than findings, and nothing here establishes what Granola did.
Granola publishes more of its chain than most: transcription through providers like Deepgram and Assembly, models from providers like OpenAI and Anthropic, notes held in a US-hosted AWS virtual private cloud, encrypted at rest and in transit. It also states that it does not store meeting audio, only transcripts and your notes.
Otter
Otter states that it trains its models on de-identified user content, and that imported files are excluded from training. On the control: Otter's Enterprise plan documents an AI model training opt-out arranged through your account manager. Its public pages do not document a self-serve toggle, so if you are on a lower tier, ask before assuming you can switch it off yourself.
Otter also carries a live legal question. In August 2025 it was named in a proposed class action in California federal court, since consolidated as In re Otter.AI Privacy Litigation, alleging that its notetaker recorded the conversations of meeting participants who were not Otter subscribers without their consent, and that recordings were used in model training. Otter denies the allegations and has moved to dismiss; the court had not ruled as of mid-2026. Allegations in a filed complaint are not findings, and the case does not establish what Otter did. It is worth knowing about because it goes to consent rather than to a feature, and because it is the clearest illustration of the point below about the person who did not choose your notetaker.
The ones that say they do not
Jamie
Jamie states that your data is never used for training of models, "not by us or any of the model providers". Jamie says it in one sentence rather than leaving you to assemble it from a subprocessor list, which is worth something. Be clear that it is a clarity win rather than a category difference, though: Fathom, Granola, Otter, Fireflies and Zoom all make an equivalent commitment about their own model providers on their own pages.
It also states that meeting audio is deleted after transcription, and that all data storage and processing stays within the EEA, Switzerland and the UK. Read that last one alongside its privacy policy, which describes an audio and summarisation pipeline running on cloud services and names a US provider handling interim data including temporary audio files under standard contractual clauses. Disclosed and lawful, but worth knowing.
Jamie holds ISO 27001 certification.
tl;dv
tl;dv's security page states plainly: no customer data is used to train the AI. It then does something unusual and describes the safeguards around its model provider. Metadata such as your name, email address and company name is anonymised before processing. Meetings are chunked into short pieces with the sequence randomised, so the provider never receives more than a fragment and cannot tell which fragments belong to the same conversation.
tl;dv states its data is hosted and stored in the EU, and any account can choose in preferences where the AI itself runs: in the US with Anthropic, or in France with Mistral. The anonymisation and chunking safeguards above describe the Anthropic path. It holds SOC 2 Type II and states GDPR and EU AI Act compliance.
Fireflies
Fireflies states that a zero-day retention policy means your meeting data is never used for AI model training, and that the same zero-day retention applies to its vendors and partners.
That is a clearer public position than several tools that market themselves harder on privacy, and it is worth saying so plainly. Fireflies is US-hosted by default, with storage in a location of your choosing available on Enterprise. It holds SOC 2 Type II and states GDPR compliance, with HIPAA and a BAA on Enterprise.
Zoom
Zoom updated its terms of service after a public argument in 2023 to state that it does not use customer content, including audio, video and chat, to train its own or third-party AI models.
This one surprises people who assume the platform vendor is the worst offender. On this specific question, Zoom's position is stronger than Fathom's, Granola's or Otter's. The reasons to look past Zoom's built-in assistant are that it only works inside Zoom, that the host controls whether it runs, and that your audio is still processed in Zoom's cloud. Training is not one of them.
The ones with nothing to train on
This is a different category and it is worth separating.
Weeve
Weeve runs the whole workflow on your Mac. It captures the machine's own audio, so no bot joins the call, and the transcript and summary are produced by a model running on the device. The recording, the transcript and the summary never leave the laptop.
There is no training policy to read because there is no copy of your meeting on anyone's server. No de-identified extract, no opt-out to find, no subprocessor list to audit. To be precise about the promise: Weeve uses online services for sign-in, billing, updates, content-free analytics and a one-off model download on first run, the same as most desktop software. What it does not do is send your meeting anywhere.
The honest trade is that Weeve is Mac only, with no Windows client and no mobile app, and there is no web app, so a recording lives on the Mac that made it.
Meetily
Meetily records, transcribes and summarises on your own machine and is open source under an MIT licence. Its onboarding installs a local summary model, so the whole pipeline is local out of the box. There is nothing to train on, and you can verify that by reading the code rather than a policy.
It also lets you point summaries at your own Claude or Groq API key if you prefer. That is a choice you make rather than a default, but it does mean the answer to "does it train on my data" depends on how you configured it.
MacWhisper
MacWhisper runs Whisper and Parakeet models locally on the Mac. On the Pro licence it also records Zoom, Teams, Webex, Skype and Discord calls in the background with no bot. Nothing is uploaded by default.
Same caveat as Meetily: Pro can route transcription to cloud services if you switch that on. Left alone, the file stays on the machine.
What "de-identified" actually means
Every vendor that trains on your meetings uses this word, and it does real work. De-identification strips the obvious identifiers: names, email addresses, account details.
What it does not do is remove the content. A de-identified transcript of a conversation about a specific merger, a specific diagnosis or a specific dispute is still a transcript of that conversation. Names are the easy part. The facts of a matter are what make it confidential, and de-identification does not touch them.
This is not an accusation of bad faith. De-identification is a legitimate and widely used technique, and the vendors doing it are being open about it. It is a reason to think about whether "de-identified" answers your particular obligation, which for a lawyer, a clinician or anyone under an NDA it often does not.
If you want to see the difference for yourself, run a real transcript through a browser-based document anonymiser. It strips names, emails, phone numbers, national insurance and IBAN numbers in the page, with nothing uploaded. What survives the redaction is the part that was actually confidential, and it is usually most of the document.
The consent problem nobody mentions
There is a second-order issue that the opt-out framing hides.
You can opt out. The person on the other side of the call cannot. They did not choose your notetaker, they may not know it is running, and they certainly did not read its privacy policy. If your tool trains on de-identified meeting content by default, their words are in the training set too, on the basis of a setting in your account.
That is the substance of the Otter complaint, and it is the reason "I turned the setting off" is a weaker answer than it sounds. It protects you. It does not describe what happens to everyone else on your calls, and it depends on you having found the setting before the meeting rather than after.
How to check any vendor yourself
The vendors change these policies, so the useful skill is checking rather than trusting a list, including this one. Four steps, in order:
Search the privacy policy for "train". Not the marketing page, not the trust badge. The policy. This finds the answer in under a minute for almost every vendor.
Check whether the statement covers model providers, not just the vendor. "We do not train our models on your data" and "your data is not used to train any models" are different sentences. The gap between them is every third-party API the product calls.
Check whether it is opt-in or opt-out, and on which plans. An opt-out that exists only on paid tiers is a different product from one available on free.
Check where the data is stored and for how long. Retention is what turns a one-off processing step into a standing archive.
If the answer to any of these takes more than a few minutes to find, that is itself information about how the vendor thinks about the question.
Which should you choose?
If your obligation is professional or regulatory: the strongest position is the one where the recording never leaves your device, because there is nothing to govern. Weeve on a Mac, or Meetily if you want open source and Windows.
If cloud is fine but training is not: Jamie, tl;dv and Fireflies all state they do not train, and Jamie's statement explicitly extends to its model providers.
If you are staying inside a platform: Zoom's position on training is clear and negative, which is more than several dedicated notetakers offer.
If you are on Fathom or Granola and want to stay: find the setting and turn it off. Both have one, on every plan, and it works. On Otter, the documented opt-out is an Enterprise control arranged through your account manager, so ask what applies to your tier. Then decide whether you are comfortable that it protects only your side of the conversation.
If it is one recording rather than a workflow: you do not need a notetaker at all. A free browser transcriber produces a transcript with speaker labels without the file leaving the page, which means no policy to read and no account to create.
The safest data is the data that was never collected.
Weeve's Starter plan is free and runs entirely on your Mac, so there is no training policy to read and nothing to opt out of. Try it on a real meeting and see whether the question needs answering at all.
If a vendor's position on training is what rules it out for you, there are fuller write-ups of the alternatives to Otter and the alternatives to Fathom, and a roundup of privacy-first notetakers in Europe for readers with an EU obligation.
FAQ
Which AI notetakers do not train on my meeting data?
Jamie, tl;dv, Fireflies and Zoom all state that they do not use customer meeting content to train AI models. Weeve, Meetily and MacWhisper do not train on it either, for a different reason: they process on your own machine by default, so the data never reaches a server.
Does Fathom train on my meetings?
Yes, unless you opt out. Its privacy policy says it may use de-identified data from meeting content to train and improve its in-house AI models, and that you can opt out in your account settings. Separately, it states that third parties such as OpenAI, Anthropic and Google are not authorised to train on your data.
Does Granola train on my meetings?
Yes by default, unless you opt out, and the opt-out is available on every plan including the free one. Enterprise workspaces are opted out by default.
Does Otter train on my meetings?
Yes, on de-identified user content. Imported files are excluded. Otter documents an opt-out on its Enterprise plan, arranged through an account manager, and does not document a self-serve toggle on its public pages. Otter is also facing a proposed class action filed in August 2025, now consolidated as In re Otter.AI Privacy Litigation, alleging recording without the consent of non-subscriber participants and use of recordings in training. Otter denies it and has moved to dismiss; those are allegations, not findings.
Does Zoom train its AI on my meetings?
No. Zoom's terms state it does not use customer content, including audio, video and chat, to train its own or third-party AI models.
Is "de-identified" good enough for confidential work?
Sometimes, and often not. De-identification removes names and identifiers, not the substance of the conversation. If your obligation is about the content of a matter rather than the identity of the parties, de-identification does not resolve it. This is a question for your own compliance position rather than a general answer.
If I opt out, is my meeting fully protected?
It stops your content being used for training. It does not change where the recording is stored, who can access it, or how long it is kept, and it does not give the other people on the call any say. Those are separate questions worth asking separately.
Do I still need consent if the notetaker runs on my own device?
Yes. Consent is about recording a person, not about where the file ends up, and in many places the law requires it. There is a fuller breakdown in the guide on whether it is legal to record a meeting.
The tools that cannot train on your meetings are the ones that never receive them.
If you would rather not read another training clause, Weeve's free Starter plan keeps the recording and the notes on a Mac, 10 a month, with nothing sent for training because nothing is sent at all. Mac only, Apple Silicon required. On any other machine the browser transcriber runs in the tab and the file stays there.


