Center for Practical AI
Education · Privacy & DataAssume it persists

The question nobody asks before typing: where does this go?

People tell AI tools things they wouldn’t put in an email — health worries, work secrets, the contents of their inbox. The tools feel private because they feel like a conversation. They aren’t. Everything you type is data held on a company’s servers, subject to its retention rules, its lawsuits, its breaches, and its next policy change. This guide is about where that data actually goes — and what you can do about it.

12 min read · Vendor practices change — every claim here is dated

Four questions

What happens to what you type.

For any AI tool, four questions decide your exposure. The answers live on the vendor's own policy pages — and they change, so this section is dated.

Is it used to train the model?Often, yes, on consumer tiers — unless you turn it off. As of mid-2026, OpenAI’s consumer ChatGPT can use your conversations to improve its models unless you change the setting; Anthropic’s September 2025 consumer terms introduced a choice to allow training with multi-year retention for those who opt in. The specific defaults move; the habit that survives every change is to check the setting on each tool you use.

Do humans read it?Sometimes. Google’s own documentation says sampled Gemini conversations may be reviewed by human reviewers, and warns users not to enter anything they wouldn’t want reviewed. “Talking to an AI” and “a person may read this later” are not mutually exclusive.

Does paying change it? Usually. Business and enterprise tiers of the major products generally state they do not train on your content by default. The single most useful sentence in this guide follows from that: use work-provided enterprise accounts for work data.The workplace reality cuts the other way when they don’t — the majority of people using unapproved “shadow” AI tools admit putting potentially sensitive data into them.

Opt-out

consumer AI often trains on your chats unless you turn it off (as of mid-2026)

OpenAI / Anthropic official policy pages

Sampled

conversations may be read by human reviewers — per the vendor's own docs

Google Gemini privacy documentation

~75%

of shadow-AI users admit entering sensitive data into unapproved tools

UpGuard 2025 — company survey

Vendor data practices change frequently. The claims in this section were verified in July 2026 against each company’s official policy pages and are on CPAI’s quarterly review schedule. Always confirm the current setting on the tool itself.

The structural fact

Your chats can outlive the chat.

Data a company holds is subject to legal process, breach, and policy change — whatever the company's intentions.

The clearest illustration came out of a copyright lawsuit. In the New York Times litigation against OpenAI, a federal magistrate ordered OpenAI in November 2025 to hand over about 20 million de-identified consumer ChatGPT conversations — a random sample spanning December 2022 to November 2024, up to 80 million prompt-and-response pairs — to opposing counsel in discovery. A request to pause the order was denied.

Read that plainly, without the outrage. Ordinary people’s chats became evidence in a lawsuit they are not party to — not because anyone did anything wrong, but because that is what happens to data a company holds. It can be subpoenaed. It can be breached. The policy that governs it today can change tomorrow. None of that requires bad intent; it is simply the nature of information that lives on someone else’s server.

Provenance

Where the training data came from.

The models were built on text and images, much of it scraped before consent norms existed. The courts are only now drawing the lines.

In the Bartz v. Anthropic case, a federal court found that training a model on lawfully purchased books can be fair use — while downloading pirated copies to do it is not. The resulting $1.5 billion settlement, granted final approval, is the largest copyright settlement in U.S. history. Both halves matter: the training itself was treated as potentially permissible; the way the data was obtained was not.

The broader point is simpler and harder to litigate. Much of the public web was scraped into training sets before anyone was asked, and as an individual you generally cannot audit whether your words or images are inside a given model. The lawsuits above are the visible edge of a provenance problem that runs underneath the whole technology.

Aggregation

Data brokers meet AI.

AI didn't create the data-broker economy. It lowered the cost of aggregating, inferring, and profiling from it.

“Anonymized” rarely means anonymous. A widely cited Nature Communications study estimated that 99.98% of Americans could be correctly re-identified in any dataset using just 15 demographic attributes. Once data can be re-linked, the distinction between anonymous and identified is mostly theoretical.

The enforcement landscape is starting to respond. The FTC’s orders against the location-data brokers X-Mode and InMarket (2024) and its 2026 settlement with Kochava — banning the sale of sensitive location data — show regulators treating precise location and other sensitive categories as off-limits to sell. AI’s role is the multiplier: it makes aggregating and profiling from broker data cheaper and faster, which is why the same personal details are worth more assembled than apart.

How scattered personal details get assembled into a profile — and weaponized — is the subject of a companion guide.

Doxxing, Disclosure, and AI →
The sensitive stuff

Health, genes, and children.

The categories where a privacy failure is hardest to undo are exactly the ones with the thinnest protections.

Genetic data. When 23andMe went bankrupt in 2025, its database of genetic profiles became a transferable asset of the estate. A court approved its sale — about $305 million — to a founder-led nonprofit, over objections from multiple state attorneys general seeking explicit customer consent. The lesson is durable: a privacy promise is only as lasting as the company that made it, and companies do not last forever.

Children’s data.The FTC’s amended COPPA Rule (effective June 2025) now requires separate verifiable parental consent before a child’s personal information is disclosed to third parties — explicitly including for AI training. It is one of the few places the law moved first.

Health and student data.Most people assume HIPAA covers anything health-related. It doesn’t — it covers providers and insurers, not wellness apps or chatbots. FERPA covers school records but has documented gaps for ed-tech. The mental-health chatbot you confide in and the learning app your child uses often sit in those gaps.

What you can do

Action for every level of influence.

1

For yourself

  • Check the training setting on every AI tool you use. Consumer versions often use your conversations to improve the model unless you turn it off. It is usually one toggle.
  • Use temporary or incognito chat modes for anything sensitive, and delete old conversations where the option exists.
  • Assume anything you type may persist — through retention, breach, legal discovery, or a change in the company's policy. Type accordingly.
2

For families

  • Never paste other people's personal information — health details, a friend's address, a child's data — into a consumer AI tool. It is not yours to hand over.
  • Review the consent flow on any AI app a child uses. New federal rules require separate parental consent before a child's data goes to third parties, including for AI training.
  • Talk about it plainly: the chatbot is not a diary. It is a company's server.
3

For working professionals

  • Use work-provided enterprise accounts for work data. Business and enterprise tiers generally state they do not train on your content by default — consumer tiers often do the opposite.
  • Do not solve the data-privacy problem by using an unapproved personal tool at work. That is how sensitive data ends up in the least-protected place.
  • If your organization has no AI data policy, ask for one before you need it.
4

For organizations & schools

  • Give people sanctioned tools with clear data terms. Where guidance is absent, people use shadow tools and put sensitive data into them anyway.
  • Map which laws actually cover your data. HIPAA and FERPA cover narrow slices; most consumer AI use falls outside them.
  • Write down what may and may not be entered into which tools — and keep it current, because vendor practices change.

Where this leads

CPAI teaches this in workshops and cohort programs.

We deliver this material to schools, libraries, employers, and community organizations — in person and online.

Sources

Research & further reading.

Official policy / primary sourceOpenAIHow your data is used to improve model performanceOpenAI's official page describing how consumer ChatGPT data may be used to improve models unless you turn the setting off, plus temporary-chat and data controls. Consumer defaults change — verify on the page and check the date below.
Official policy / primary sourceAnthropic (2025)Updates to Consumer Terms and Privacy PolicyAnthropic's September 2025 consumer-terms update introduced a choice to allow chats to be used for model training, with extended (multi-year) retention for those who opt in. Official announcement; verify current defaults at the link.
Official policy / primary sourceGoogleGemini Apps privacy notice & data controlsGoogle's own documentation states sampled Gemini conversations may be reviewed by human reviewers and warns users not to enter information they wouldn't want a reviewer to see or used to improve services. Official support page; verify defaults at the link.
Company-reportedUpGuard (2025)The State of Shadow AICompany-reported survey: 81% of employees use unapproved AI tools, and conventional awareness training did not reduce risky use. About three-quarters of shadow-AI users admit entering potentially sensitive data into unapproved tools.
Survey / industry reportUniversity of Melbourne & KPMG (2025)Trust, Attitudes and Use of AI: A Global StudyAcademic-led global survey, n≈48,000 across 47 countries — the strongest of the survey tier. Only 47% of workers report any AI training; 66% rely on AI output without evaluating accuracy; 56% report AI has caused mistakes at work; 57% hide their AI use from employers.
Court ruling / legal filingNYT v. OpenAI (2025)Court order to produce 20 million ChatGPT logsIn the New York Times copyright litigation, a federal magistrate ordered OpenAI (November 2025) to produce about 20 million de-identified consumer ChatGPT conversations — a random sample from Dec 2022–Nov 2024, up to 80 million prompt-output pairs — to opposing counsel in discovery; a stay was denied. Ordinary users' chats became evidence in a case they are not party to. A structural fact about data held by a company, not a scandal claim.
Court ruling / legal filingBartz v. Anthropic (2025)Fair-use ruling and $1.5B copyright settlementA federal court found that training on lawfully purchased books can be fair use, while downloading pirated copies is not — and the resulting $1.5 billion settlement (final approval granted) is the largest copyright settlement in U.S. history. State both halves: the fair-use holding and the piracy liability.
Peer-reviewed studyRocher, Hendrickx & de Montjoye (2019)Estimating the success of re-identifications in incomplete datasetsNature Communications. A model estimated that 99.98% of Americans could be correctly re-identified in any dataset using 15 demographic attributes — so 'anonymized' data is often re-identifiable once combined with other data.
Official policy / primary sourceFederal Trade Commission (2024–2026)Enforcement against location-data brokersThe FTC's orders against X-Mode/InMarket (2024) and its settlement with Kochava (May 2026, banning the sale of sensitive location data) illustrate the enforcement landscape for data brokers that aggregate and sell precise location information.
Official policy / primary sourceFederal Trade Commission (2025)Amended COPPA RuleThe FTC's amended Children's Online Privacy Protection Rule (final April 2025, effective June 23, 2025) requires separate verifiable parental consent before a child's personal information is disclosed to third parties — including for AI training.
Court ruling / legal filingIn re 23andMe (2025)Bankruptcy and sale of genetic dataAfter 23andMe filed for bankruptcy in 2025, its genetic database became a transferable estate asset; a court approved its sale (about $305 million) to TTAM Research Institute, a nonprofit led by the company's founder, over objections from multiple state attorneys general seeking explicit customer consent. The teaching point: a privacy promise is only as durable as the entity that made it.
Official policy / primary sourceIAPP / MultiState (2026)U.S. state comprehensive privacy lawsThere is no comprehensive federal consumer-privacy statute; sector rules (HIPAA, FERPA, COPPA, GLBA) cover slices. About 20 states have comprehensive consumer-privacy laws in effect as of 2026 — North Carolina is not among them.
Official policy / primary sourceCooley / state trackers (2026)State AI laws — current landscapeA brief factual map: Colorado's AI Act (amended, effective date delayed to 2026), Texas's TRAIGA (effective Jan 1, 2026), Illinois's restriction on AI-delivered therapy, and California's SB 243 for companion chatbots. A maintained secondary tracker keeps the current picture. Described, not scored.
Last reviewed: July 2026We review this page quarterly. Statistics in this category change rapidly.Vendor data practices (Section 1) change frequently and are verified against each company's official policy pages; confirm the current setting on the tool itself. The legal landscape moves — the state-law count and North Carolina's status are re-checked each quarter. Litigation described here (NYT v. OpenAI, Bartz v. Anthropic, the 23andMe sale) was current as of July 2026.

Want CPAI to teach AI data safety in your school or workplace?

We deliver this material as workshops and cohort programs for schools, libraries, employers, and community organizations.