Privacy primer // September 2026

How solving Navier–Stokes turned into a privacy debate

Apparently it took a math controversy to get capability researchers talking about privacy. Let’s go through what the labs actually do with our data.

These policies change often. Everything below reflects the labs' published policies as I read them on September 9, 2026; the links point to the pages I checked.

I’ve been working on privacy in AI for years, and I’ve been beating the drum about how murky and misleading frontier labs’ privacy and data-retention policies are for just as long. Here’s a tweet I wrote about Anthropic’s policies almost two years ago. In my experience, people usually gasp, get concerned for a moment, and then move on. Some of the people working on model capabilities dismiss these concerns as “non-technical.”

So it’s interesting to me how we ended up here, with capability researchers getting involved and posting opinions about privacy. It all started with OpenAI announcing what it described as a solution to the Navier-Stokes Millennium Prize problem. Levent Alpöge and Tristan Buckmaster had been working for much of the past year on related problems. Buckmaster questioned why OpenAI had pursued a similar approach after hearing about their progress. OpenAI said its effort began after hearing those rumors, and that it used roughly 10,000 agents. Buckmaster’s account, OpenAI’s account

Then things got to privacy and data retention when OpenAI initially wrote:

“While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

OpenAI’s statement

That raised alarms about whether OpenAI had gone through their logs. OpenAI denied that its researchers or agents had accessed their work before it was public. In a September 10 update, it went further: after an investigation, it said Buckmaster’s Codex prompts from the two months before the announcement could not have influenced the system, including through training. That is an important update to the original statement. OpenAI’s updated account

But I think focusing only on whether someone read the chat logs misses something. Sometimes the valuable secret is just one bit: someone has already made this approach work with an AI. You don’t need their proof, their name, or their conversations for that information to change where you put your time and compute. OpenAI’s own account says hearing rumors of progress prompted its effort. That illustrates the value of the signal; it does not establish that the signal came from private chats.

This is why I care about worst-case privacy, not just average-case performance. A system can scrub identifying details from almost every conversation and still expose the one fact that matters most to a particular person. “We removed the names” doesn’t answer that concern. Neither does a low average leakage rate.

And this is why I don’t think Clio-style aggregation and LLM privacy filters are enough. Clio is designed to suppress rare patterns, but its published safeguards do not provide an end-to-end differential privacy guarantee. CLIOPATRA shows a concrete failure: researchers inserted malicious chats into a local reproduction of Clio and recovered sensitive medical information from the resulting aggregate summaries. Those tests used synthetic target conversations, not Anthropic’s live service. I’m not saying Clio was involved in the Navier–Stokes episode. I’m saying that hiding raw chats and checking summaries with another model does not settle what someone can learn. More on these systems below.

So what is the truth? What do we know about the privacy and data-retention policies of frontier AI labs? Let’s go through them!

1. What do the basic policies say for different tiers?

First, paying for a subscription does not automatically get you different privacy terms. ChatGPT Plus is a personal subscription. Claude Pro is a personal subscription. Those are different products from an employer’s managed business account, even if you’re talking to the same model.

There are also two separate things to check: can they use your conversations to train models, and how long do they keep them? Turning off training doesn’t necessarily delete anything. It can just mean the conversation stays in your history without being used for that purpose.

OpenAI: On personal ChatGPT accounts, including paid subscriptions, conversations can be used for training unless you opt out. Ordinary chats stay in your account until you delete them. The familiar “30 days” refers to the deletion period after you delete a chat, subject to exceptions. It does not mean everything you said last month has already disappeared. Training policy, retention policy

Anthropic: Claude Free, Pro, and Max have a model-improvement setting. If you allow model improvement, new or resumed chats can enter training pipelines, where de-identified data may stay for up to five years. If you delete a conversation, Anthropic says it removes it from backend storage within 30 days, with exceptions. Again, that is a deletion timeline, not a promise that all your undeleted chats expire after a month. Training policy, retention policy

Google: For adults using Gemini with a personal account, Keep Activity is on by default. Saved activity can be used to improve models. The default auto-deletion period is 18 months, which you can change. Turning Keep Activity off stops future chats being used for training unless you submit feedback, but Google still keeps those conversations for up to 72 hours. Paying for a personal Google AI plan does not give you the same terms as a qualifying Workspace account. Activity controls, consumer privacy policy

Grok: Conversations on Grok.com and its mobile apps can be used for training unless you opt out. Its FAQ says new conversations are excluded once you turn that off. Deleting a conversation generally starts a deletion period of up to 30 days, again with exceptions. Grok on X has separate controls, so don’t assume changing a setting in one place changes it everywhere. Grok’s FAQ, Grok on X

Business accounts and APIs often offer stronger protections. OpenAI’s business products and API, Anthropic’s commercial products, qualifying Google Workspace accounts, and Grok’s API generally exclude customer content from model training by default or require permission. Those protections matter. But “we don’t train on it” still doesn’t mean “we don’t keep it.” For example, Grok’s API ordinarily keeps requests and responses for 30 days for abuse auditing. OpenAI, Anthropic, Google Workspace, Grok API

This is where zero data retention, or ZDR, comes in. It is a specific arrangement about what the provider stores, for particular services and features. It isn’t something you should assume comes with every enterprise account. For example, OpenAI requires approval for its API ZDR controls, and some features can still store data needed to provide the service. You have to check the actual coverage. OpenAI’s API data controls

2. What are the carve-outs and exceptions you should watch out for?

This is the part that tends to surprise people. You can understand the ordinary policy, choose a privacy setting, and still miss a separate rule that applies to your conversation.

2.1 Feedback

Say you’ve turned off training. You spend an hour discussing unpublished work, then click thumbs-down because the last answer was wrong. Did you just rate that answer, or did you agree to let the company use the conversation?

OpenAI says feedback can make the entire associated conversation eligible for training, even if you previously opted out. Yes, the whole conversation associated with that feedback, not just the response you rated. OpenAI’s feedback policy

Anthropic also includes the entire related conversation, with retention for up to five years. Its policy excludes raw material from connectors unless that material is directly copied into the conversation. So this does not automatically mean every file in a connected account is included, but it can include much more than the one answer you meant to comment on. Anthropic’s feedback policy

Google’s consumer Gemini policy says that feedback submitted with Keep Activity off can include the previous 24 hours of chats, including uploads and Connected App content. Grok on X also makes an exception for feedback and the associated conversation. Gemini’s feedback policy, Grok on X

This is what I find misleading about these interfaces. You’ve made a choice about training, and then a small feedback button can change what the company is allowed to do with a conversation. If that’s the agreement, say so next to the button. People shouldn’t have to infer it.

2.2 Safety flags

Another reason a provider may keep a conversation is that an automated system flags it for a possible policy violation. The question then becomes what the flag allows: longer storage, human review, or some other use.

Anthropic says flagged inputs and outputs can be kept for up to two years, and trust and safety classification scores for up to seven years. Those are different records. The seven-year period is for the scores, not a blanket seven-year period for full conversations. Anthropic’s retention policy

OpenAI can also tell certain API customers that flagged conversations will be kept and reviewed, including where ZDR previously applied. Its Safety Retention policy requires advance written notice and applies to particular customers and models when needed to investigate or prevent severe risk activity. Images flagged as possible child sexual abuse material have a separate review exception even with ZDR enabled. OpenAI’s API data controls

So a flag can have consequences for your privacy. That makes it reasonable to ask how the system handles mistakes, who gets access to the conversation, and whether the extra copy will actually be deleted. Calling something a safety classifier does not answer those questions.

And the Fable situation has already changed. In June, Anthropic introduced 30-day retention for Fable 5 and other designated models, affecting uses previously covered by ZDR. That created a problem for businesses that could not accept the provider storing their conversations. Anthropic acknowledged this in September and announced Enterprise Frontier Safeguards, with options for customers to keep monitoring data in their own infrastructure and have their own staff review alerts. Covered-model policy, September announcement

While Anthropic gets that system ready for a rollout later this fall, eligible customers can temporarily use Fable 5 and 5.1 with ZDR for internal business applications. Real-time classifiers still operate. So a request can be checked for misuse without the ordinary 30-day storage requirement applying. Whether your data is retained depends on the model, agreement, and features you use. Transition terms

2.3 Legal proceedings

Even if a company wants to delete your data, a court may require it to keep it.

We’ve already seen this in the New York Times litigation against OpenAI. A 2025 order required OpenAI to preserve data it would otherwise delete. OpenAI says the broad obligation ended on September 26, 2025; its October update described continued preservation of some historical data. So that particular broad order is no longer in effect. OpenAI’s update

The part that should concern users is that they didn’t have to be parties to the lawsuit for their conversations to be affected. Their deletion preferences could still be overridden. The order also had exclusions, including qualifying ZDR API use. Whether the provider has a record to preserve makes a real difference. Scope of the order

2.4 Aggregate statistics and reviews of how people use the models

Another use to watch for is analyzing conversations to understand how people use the product. A company can do this without training the model you talk to. Google, for example, says its Gemini settings do not control processing chats to create anonymized data for service improvement. So a training opt-out does not answer every question about analysis. Gemini’s policy

Anthropic is public about its approach. Clio, now called Anthropic Insights, summarizes and groups conversations, with minimum group sizes and automated privacy checks. Anthropic says employees see aggregate findings through its research mode without access to raw conversations. It separately describes a safety mode whose results can be linked to accounts and reviewed by authorized staff. These are different uses, with different protections. Clio’s privacy explanation, research announcement

But “aggregate” does not settle the privacy question. Clio’s published design relies on models following instructions to remove sensitive information and other models checking the results. It does not provide an end-to-end differential privacy guarantee. The original system used Claude 3 Haiku to summarize individual conversations, with stronger models downstream. We need to test the models doing the scrubbing, including the smaller ones. Clio paper

And the concern is already empirical. CLIOPATRA demonstrated targeted extraction of medical information from a local reproduction of Clio, including a configuration using Haiku. An attacker with some information about a target inserted malicious chats and inspected the resulting aggregate summaries. The tests used synthetic medical conversations, not private chats from the live service. Still, they show that grouping conversations and running privacy filters does not reliably prevent attribute inference.

There is also a very basic problem when an LLM is asked to sanitize another LLM query: it can confuse the query it is supposed to scrub with an instruction it should follow. Say the outer instruction is to remove names, but the text inside asks to preserve a name. The model can obey the inner request and leave the name in. This can happen with ordinary instruction-like text, without someone deliberately attacking the system.

We tested this in Can Large Language Models Really Recognize Your Name?, using Clio’s published prompts with Claude 3.5 Haiku for summarization and Claude 3.7 Sonnet for auditing. Adding instruction-like text increased ambiguous-name leakage from 3.92% to 18.20% in these component tests. The auditor missed some leaks too. Those are controlled test results, not a measured leakage rate for deployed Clio.

I would be even more cautious about extending these protections to heavily heterogeneous, multimodal data: text, images, voice, and files can each supply another identifying detail. The papers above demonstrate failures with text. They do not test that whole setting. A privacy claim covering it needs evidence that the system handles combinations of details, including what someone can infer without seeing a name.

There is work on stronger alternatives. Urania builds aggregate summaries with differential privacy, which bounds how much one conversation can affect what is released. Its guarantee is per conversation, so protecting someone who contributes many chats needs additional work. That is the kind of explicit guarantee and limitation I want to see, instead of treating an LLM’s privacy check as the end of the discussion.

2.5 There are hidden exemptions to opt-outs

The wording around training opt-outs deserves its own discussion. OpenAI’s Data Controls FAQ says the setting stops conversations from training “ChatGPT.” Its broader explanation says new conversations won’t train “our models.” It also offers a privacy portal. These policies do not establish that the portal is a required second opt-out, or that internal reward models are exempt. But why should users have to compare pages to figure out the scope of one setting?

There is a documented separate control: training on full Codex environments has its own settings, which neither the ChatGPT toggle nor the privacy portal changes. That is a concrete example of why the product and type of data matter. OpenAI’s data-use explanation

And the distinction between models matters too. A preference for one answer over another can train a reward model, which scores responses and helps steer the model you eventually talk to. OpenAI describes this in its explanation of ChatGPT’s training. It has also explicitly said that a GPT-4o update added a reward signal based on users’ thumbs-up and thumbs-down feedback. So feedback can change how future models behave. OpenAI’s account of the sycophancy update

So I want the labs to spell out whether an opt-out covers the model answering me, reward models, safety classifiers, and other internal models. What happens to preferences inferred from my interactions, or examples used for evaluation? Which analysis continues after I opt out? OpenAI’s privacy policy separately describes research and uses of aggregated or de-identified information. Users deserve a clear explanation of these uses.

3. Can you delete your chats? Really?

You can delete them from your history. What happens beyond that depends on which copies exist and which exceptions apply.

For example, OpenAI’s deletion policy has an exception for conversations already de-identified and disassociated from you, as well as legal and security exceptions. Google says deleting Gemini activity does not delete conversations and related data already reviewed by humans; those records can be retained for up to three years. So “I deleted it” doesn’t necessarily mean “the company no longer has any version of it.” OpenAI’s deletion policy, Google’s review exception

The information can also exist in more than one place. You might upload a document, discuss it in a chat, and have information from it added to memory. OpenAI manages files saved to Library separately from chats, and its memory guidance says to remove information from the sources where it appears. Deleting the conversation alone may leave the file or remembered information available. File retention, memory controls

And temporary modes are not the same as immediate deletion. ChatGPT’s Temporary Chat can retain a copy for up to 30 days for safety; Gemini temporary chats remain for 72 hours; Grok’s Private Chat generally uses a 30-day deletion window. These modes offer useful protections, but the names don’t tell you the retention period. ChatGPT, Gemini, Grok

There are meaningful controls, too. Anthropic says deleting a chat excludes it from future training, and turning off model improvement excludes previous and new chats from future training runs. But neither action reverses training already completed or underway. Deleting a record and undoing what a model learned from it are different problems. Anthropic’s training controls

Basically, use the deletion controls, but check what they delete. A conversation, a saved file, a memory, a feedback submission, and a safety record may all be subject to different rules.

4. What does using your data actually mean?

Going back to the discussion that started this: “training on user data” can mean several things. A model might be trained to predict the text you wrote. Your prompts might be paired with another model’s answers to train a smaller model. Your work might be turned into tasks or reward signals for reinforcement learning. These uses don’t have identical privacy or intellectual-property implications. The training distinctions discussed in the thread

For example, using a private coding task to build a training exercise raises a question about repurposing that work, even if the resulting model never repeats the code word for word. The feedback and aggregate-analysis examples above show why we need to ask about the whole process, including the information used to evaluate and steer models.

There is also a reason to be careful with broad claims that RL avoids privacy problems. In a recent preprint, we found that RL on unrelated factual questions increased extraction of email addresses the tested models had already memorized. That is about making existing memorized information easier to extract, rather than showing that RL memorized new private conversations. But it should make us cautious about treating the training method as a privacy guarantee. The RL study

5. What about de-identified and synthetic data?

This is another phrase that gets used as if it settles the issue. It doesn’t.

You can remove someone’s name and email address and still leave enough information to identify them. Think about a conversation describing an unusual job, a particular city, and a very specific event. Those details may be enough when combined with information from elsewhere. They can also reveal something sensitive about a person without ever giving you their name.

We tested this in A False Sense of Privacy. Using additional information, attackers could match records and infer sensitive attributes from text that had been sanitized, including through synthetic generation. The experiments were on released datasets. They show why removing obvious identifiers is not enough; they don’t establish that a particular provider’s current system has been compromised.

Generating synthetic text from a real conversation doesn’t automatically remove the information that made the original sensitive. We should be asking how a method was tested and what it protects against.

This is also the point of Privacy Is Not Just Memorization. A provider can retain a sensitive conversation, another system can receive it through a connected tool, or someone can infer private information from it. None of that requires a model to regurgitate a training example. A low risk of regurgitation is good news about one risk. It does not resolve the others.

6. Can retention and monitoring be useful for safety?

Yes, and I don’t think we should dismiss that. There are cases where monitoring and limited collection are necessary for the safety function we want a system to perform. If we expect it to recognize escalating self-harm risk, it needs to process enough of the conversation to notice what is happening. Evaluating whether it responds appropriately also requires evidence, including carefully collected and labeled examples. This matters especially for teens and people using these systems for mental health support.

There are already concrete examples. Anthropic describes a self-harm classifier that triggers support resources, and safety evaluations using voluntarily shared conversations. OpenAI’s parental controls allow limited safety notifications after trained reviewers identify serious self-harm concerns, without giving parents access to the teen’s chat history. These are examples of monitoring and review being used for a specific purpose; they do not establish that keeping everyone’s conversations for years prevents harm. Anthropic’s wellbeing safeguards, OpenAI’s parental controls

What I want to see is privacy-preserving research on monitoring streams of interactions. Can we process a short window of conversation, record the risk label and when it occurred, and then discard the underlying messages? Could coarse timestamps or elapsed time tell us enough about escalation? For research on whether an intervention helped, could we retain a small set of labels and outcomes, with properly obtained consent, instead of a permanent transcript?

These are research questions, not guarantees that those records would be harmless. A self-harm label is sensitive, and a timestamp can help link it to a person. We need to test how little information is sufficient, protect the remaining records, and measure missed crises as well as false alarms. Aggregate research needs privacy guarantees across repeated contributions from the same person. An individual safety alert needs its own narrowly defined access and disclosure rules. When reviewing the original text is necessary, that access and retention should be limited to the purpose.

The changes around Fable are relevant here: the location of stored monitoring data and who reviews alerts are choices the system can be designed around. I want more work on those choices, with clinicians, teens, and people who use these systems for support involved in deciding what help should look like and what information it requires.

The actual protections are much more specific than people assume. Check your account’s training setting, read what feedback does, and look at the exceptions to deletion. For confidential work, check the agreement for the model and features you actually use.

And the labs should make this easier. If a thumbs-down changes permission to use a conversation, say that at the point of feedback. If a safety flag changes retention, explain the consequences and how mistakes are handled. If deleting a chat leaves other records behind, tell the user what remains. These are concrete questions about how the systems work. Calling them non-technical has never made them go away.

Acknowledgments

Thanks to Andrew, Atoosa, and Neal Mangaokar for their feedback and for the discussions that shaped this post. This post was written with the help of Astra; the reading of the policies, the opinions, and any mistakes are mine.

← Back to writing