The answer is never the brand. It is the product, the plan, and increasingly the place you run the model. The same OpenAI model carries different contractual terms depending on whether you call OpenAI or Microsoft. Anthropic will not train on your commercial account until someone clicks thumbs up. Mistral's zero data retention setting does not apply to its preview models. These are not loopholes — they are written into the terms, and this is what they say.
Someone asks a reasonable question — if I connect an agent to my email, does the provider keep it? — and gets four different answers from four blog posts, none of which say which plan they mean.
Training policy is set by product, plan and deployment, not by company. Anthropic will not train on Claude for Work conversations and may train on Claude Pro conversations. OpenAI has not trained on API data since March 2023, while its consumer app is governed separately. Run an OpenAI model through Microsoft Foundry and Microsoft's terms apply instead of OpenAI's.
So "is provider X safe" is unanswerable. The question that has an answer is: which product, on which plan, in which region, under which agreement.
Four different things get called "training"
Most comparisons collapse these. They have different triggers and different consequences.
Training — is your content used to improve the model? Usually the clearest contractual answer.
Retention — how long is it stored, and can that be reduced? A provider can be barred from training while still holding your data for years.
Human review — can a person read it? Usually triggered by a safety flag, and frequently confused with training.
Abuse monitoring — a separate pipeline with its own clock and its own storage, usually not optional.
A "we don't train on your data" headline addresses the first and tells you nothing about the other three.
What the terms say
Read from each provider's own documentation on 20 August 2026.
| Product | Trains by default? | Retention | Notable exception |
|---|---|---|---|
| OpenAI business (Business, Enterprise, Edu, API) | No | Abuse logs ≤30 days; stateful endpoints "until deleted" | Zero retention needs OpenAI's prior approval |
| OpenAI consumer (Free, Plus, Pro, Go) | Yes — unless you opt out | Per your data controls | Opt-out applies to new conversations only |
| Anthropic commercial | No | 30 days standard | Feedback click → training + 5-year retention |
| Anthropic consumer (Free/Pro/Max) | Only if you allow it | 5 years if allowed, else 30 days | Safety-flagged chats used regardless |
| Mistral (commercial, eff. 5 Aug 2026) | Depends on the product's default | Not stated in terms | Feedback, Labs and Preview models are all carved out |
| Google Gemini (consumer) | Subset human-reviewed | 18 months default | Reviewed chats kept 3 years, separate clock |
| Microsoft Foundry / Azure OpenAI | No — and not shared with OpenAI | Customer-controlled | "Global" deployments process in any geography |
| Microsoft 365 Copilot | No — barred contractually | Per your M365 policy | Anthropic models excluded from EU Data Boundary |
| AWS Bedrock | Model providers have no access | Customer-controlled | — |
Four things in that table deserve more than a row.
Same company, opposite defaults
OpenAI is the clearest illustration that the brand tells you nothing.
Business tiers — ChatGPT Business, Enterprise, Edu, and the API since 1 March 2023 — are not trained on by default. Consumer tiers — Free, Plus, Pro and Go — are the reverse: conversations are used to improve models unless the user opts out, under Settings → Data Controls → "Improve the model for everyone."
Same company, same underlying models, opposite defaults. Which one an employee happens to be signed into decides the answer, and nothing in the interface announces it.
Two practical consequences. Opting out on a consumer account applies to new conversations only — it does not reach back over what has already been sent. And an employee who uses a personal Plus account for work is on the training-by- default side of that line, whatever your company policy says.
The feedback button is the pattern, not the exception
Three independent sets of commercial terms carve out an exception for user feedback — Anthropic's, Mistral's and Microsoft's — which makes it a category norm rather than one vendor's quirk.
Anthropic, commercial products: "By default, we will not use your inputs or outputs from our commercial products to train our models." Then — if a user "explicitly report[s] feedback or bugs to us (e.g. via our thumbs up/down feedback button)," Anthropic "may use your chats and coding sessions to train our models" and will "store the entire related conversation… in our secured back-end for up to five years."
Mistral, commercial terms effective 5 August 2026, section 4.2: Mistral will not use Customer Data or Outputs for training except — among other carve-outs — "(b) when Customer or an End User provides Feedback to Mistral AI."
Microsoft, for Copilot: "We may use customer feedback, which is optional, to improve Microsoft Copilot" — though Microsoft is explicit that this feedback is not used to train the foundation models, which is a meaningfully narrower carve-out than the other two.
Read that operationally. Your account is protected by contract. An employee finds a reply unhelpful, clicks thumbs down out of habit, and on two of these three providers that entire conversation — not the rating, the conversation — leaves the protected default.
Nobody has done anything wrong. The button is right there and it looks like product feedback. But if that conversation contained a client's contract terms, the terms governing those contract terms just changed, and no administrator was consulted.
If you run agents on any of these, this belongs in team onboarding, not in a policy document nobody opens.
An employee clicks thumbs down out of habit, and that entire conversation enters a five-year retention window.
"Zero data retention" does not always mean zero
Two providers document limits on their own strongest privacy setting.
Mistral, section 4.3: "By using Labs or Preview Models, you acknowledge that (i) Mistral AI may use Customer Data and Outputs generated from Labs or Preview Models to train its artificial intelligence models and (ii) the opt-out preferences you selected for other Mistral AI Products (including through zero data retention) does not apply to Labs or Preview Models."
Zero data retention, explicitly disapplied — for exactly the newest models a team is most likely to be trialling.
OpenAI is narrower but similar in shape. Zero Data Retention is endpoint-specific and requires "prior approval by OpenAI." Eligible endpoints include /v1/chat/completions, /v1/responses and /v1/embeddings. Stateful features — Assistants, threads, vector stores — "may still store application state, even if Zero Data Retention is enabled." Turning ZDR on does not make a vector store forget.
Deleting a chat does not always delete it
Google's consumer Gemini position has more layers than yes or no.
A subset of chats are "reviewed by human reviewers (including Google's trained service providers) to help improve Google services." You can stop future chats being reviewed by turning off Keep Activity. Default retention is 18 months, adjustable to 3 or 36.
The part that surprises people: chats already selected for human review are retained for up to three years — a separate clock that survives the deletion you just performed. Turning Keep Activity off works forwards, not backwards. Temporary chats and chats with Keep Activity off are held 72 hours.
Where you run the model changes the contract
Where a model runs decides which contract governs your prompts, and for European businesses that is the distinction that matters most. Most comparisons miss it entirely.
The same model, different terms. Microsoft Foundry states that your prompts, completions, embeddings and training data "are NOT available to OpenAI or other providers of Models sold by Azure" and "are NOT used by providers… to improve their models or services." It goes further: models sold by Azure "do NOT interact with any services operated by providers… for example, OpenAI (e.g. ChatGPT, or the OpenAI API)." An OpenAI model reached through Azure is governed by Microsoft's agreement, not OpenAI's.
AWS Bedrock takes the same architectural approach: model providers have no access to the deployment accounts, so "they don't have access to Amazon Bedrock logs or to customer prompts and completions."
But the deployment type moves your data. Azure distinguishes Global, DataZone and standard deployments. For "any deployment type labeled 'Global,' prompts and responses may be processed in any geography where the relevant model… is deployed." DataZone confines processing to the specified zone — an EU-resource DataZone deployment stays within EU member nations. Standard stays in your geography.
Choosing "Global" for capacity reasons is a data-residency decision whether or not anyone framed it that way.
One further trap for EU buyers: Microsoft 365 Copilot is an EU Data Boundary service, but its documentation states "models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary." Selecting a model inside a compliant product can move data outside the boundary that made it compliant.
There is a genuine bright spot: for Azure deployments in the EEA, Microsoft states the authorised human reviewers for abuse monitoring "are located in the European Economic Area."
If the data belongs to your clients
For an agency this stops being preference and becomes disclosure. Three things to confirm before client data flows anywhere:
Which tier the agent runs on. Commercial and consumer terms diverge most exactly where client data is most sensitive. An agent built quickly on someone's personal account is not covered by the commercial terms you would cite if asked.
Which region, and which deployment type. "We use Azure" is not a data-residency answer. Global, DataZone and standard are three different answers.
What your own contracts already promise. If you have told a client their data stays in the EU, model selection inside your tooling is now a contractual matter, not a technical preference.
The account someone is signed into decides the answer. An employee using a personal Plus account for work sits under consumer terms, where training is the default, whatever your company policy says.
How to check your own setup
1. Name the product and plan for every AI tool touching business data. Not "we use ChatGPT" — ChatGPT Team, or the OpenAI API on account X, or GPT-4o via Azure Foundry, DataZone, EU. The plan and the path are the answer.
2. Find the feedback and model-improvement controls and decide deliberately. On consumer tiers it is a setting. On commercial tiers the feedback button is the exposure. Someone should have chosen rather than inherited.
3. Check whether you are on preview or Labs models anywhere. Mistral disapplies zero data retention for these. Assume other providers may treat previews differently and verify rather than infer.
4. Read your retention clock, then look for the second one. Safety review, abuse monitoring and feedback pipelines each run their own, and those usually survive the delete button.
What this does not tell you
Everything above is a reading of published terms on a single date — not legal advice, and not a substitute for a DPA. Terms move: Mistral's commercial terms took effect fifteen days before this was written.
One sourcing caveat, stated plainly. OpenAI's own policy pages could not be opened directly — openai.com/policies, openai.com/enterprise-privacy and the help centre all returned 403. The OpenAI consumer and business rows above come from search results citing those same pages rather than from the pages themselves. The distinction is real and worth knowing: every other figure in this article was read from the source document. Verify the OpenAI rows against the live pages before relying on them for a contractual decision.
Cohere's data usage policy returned 404 and is absent entirely.
And negotiated enterprise agreements routinely override published terms. If you have one, that document wins over anything here.
Frequently asked questions
Does ChatGPT train on your data? It depends entirely on the tier. ChatGPT Business, Enterprise, Edu and the API are not used for training by default — the API since 1 March 2023. Consumer tiers, meaning Free, Plus, Pro and Go, are the reverse: conversations are used to improve models unless the user opts out under Settings, Data Controls, Improve the model for everyone. Opting out applies to new conversations only and does not reach back over what was already sent.
Does ChatGPT Business train on your data? No, not by default. OpenAI's business tiers — ChatGPT Business, Enterprise, Edu and the API — are not trained on unless the customer explicitly opts in to share data. The distinction that matters is the account someone is signed into: an employee using a personal Plus account for work sits under consumer terms, where training is the default, regardless of company policy.
Does Claude train on your data? Anthropic's commercial products — Claude for Work, the API and Claude Gov — are not used for training by default. There is one significant exception: if a user reports feedback through the thumbs up or thumbs down button, Anthropic may use those chats for training and stores the entire related conversation for up to five years. Consumer Claude tiers train only if the user allows it, and allowing it extends retention to five years rather than the standard thirty days.
Is my data safe if I use AI through Microsoft or AWS instead? The contract changes with the path. Microsoft states that prompts and completions for models sold by Azure are not available to OpenAI or other model providers and are not used to improve their models, and that such models do not interact with services operated by those providers. AWS Bedrock is architecturally similar: model providers have no access to the deployment accounts, and therefore no access to logs, prompts or completions. The same underlying model can carry different terms depending on who you buy it through.
Sources
- OpenAI, Your data (API),
developers.openai.com/api/docs/guides/your-data— retrieved 20 Aug 2026 - OpenAI, How your data is used to improve model performance and Data Controls FAQ,
help.openai.com— accessed via search results, not opened directly (403), 20 Aug 2026 - Anthropic, Is my data used for model training? (commercial),
privacy.claude.com/en/articles/7996868— retrieved 20 Aug 2026 - Anthropic, Is my data used for model training? (consumer),
privacy.claude.com/en/articles/10023580, dated 16 Mar 2026 — retrieved 20 Aug 2026 - Mistral, Commercial Terms of Service,
legal.mistral.ai/terms/commercial-terms-of-service, effective 5 Aug 2026 — retrieved 20 Aug 2026 - Google, Gemini Apps privacy,
support.google.com/gemini/answer/13594961; Privacy Notice updated 29 Jun 2026 — retrieved 20 Aug 2026 - Microsoft, Data, privacy and security for Foundry Models sold by Azure,
learn.microsoft.com/.../responsible-ai/openai/data-privacy, updated 5 Jun 2026 — retrieved 20 Aug 2026 - Microsoft, Data, Privacy and Security for Microsoft Copilot,
learn.microsoft.com/.../microsoft-365-copilot-privacy, updated 18 Aug 2026 — retrieved 20 Aug 2026 - AWS, Data protection in Amazon Bedrock,
docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html— retrieved 20 Aug 2026