The question nobody asks before typing: where does this go?
People tell AI tools things they wouldn’t put in an email — health worries, work secrets, the contents of their inbox. The tools feel private because they feel like a conversation. They aren’t. Everything you type is data held on a company’s servers, subject to its retention rules, its lawsuits, its breaches, and its next policy change. This guide is about where that data actually goes — and what you can do about it.
12 min read · Vendor practices change — every claim here is dated
What happens to what you type.
For any AI tool, four questions decide your exposure. The answers live on the vendor's own policy pages — and they change, so this section is dated.
Is it used to train the model?Often, yes, on consumer tiers — unless you turn it off. As of mid-2026, OpenAI’s consumer ChatGPT can use your conversations to improve its models unless you change the setting; Anthropic’s September 2025 consumer terms introduced a choice to allow training with multi-year retention for those who opt in. The specific defaults move; the habit that survives every change is to check the setting on each tool you use.
Do humans read it?Sometimes. Google’s own documentation says sampled Gemini conversations may be reviewed by human reviewers, and warns users not to enter anything they wouldn’t want reviewed. “Talking to an AI” and “a person may read this later” are not mutually exclusive.
Does paying change it? Usually. Business and enterprise tiers of the major products generally state they do not train on your content by default. The single most useful sentence in this guide follows from that: use work-provided enterprise accounts for work data.The workplace reality cuts the other way when they don’t — the majority of people using unapproved “shadow” AI tools admit putting potentially sensitive data into them.
consumer AI often trains on your chats unless you turn it off (as of mid-2026)
OpenAI / Anthropic official policy pages
conversations may be read by human reviewers — per the vendor's own docs
Google Gemini privacy documentation
of shadow-AI users admit entering sensitive data into unapproved tools
UpGuard 2025 — company survey
Vendor data practices change frequently. The claims in this section were verified in July 2026 against each company’s official policy pages and are on CPAI’s quarterly review schedule. Always confirm the current setting on the tool itself.
Your chats can outlive the chat.
Data a company holds is subject to legal process, breach, and policy change — whatever the company's intentions.
The clearest illustration came out of a copyright lawsuit. In the New York Times litigation against OpenAI, a federal magistrate ordered OpenAI in November 2025 to hand over about 20 million de-identified consumer ChatGPT conversations — a random sample spanning December 2022 to November 2024, up to 80 million prompt-and-response pairs — to opposing counsel in discovery. A request to pause the order was denied.
Read that plainly, without the outrage. Ordinary people’s chats became evidence in a lawsuit they are not party to — not because anyone did anything wrong, but because that is what happens to data a company holds. It can be subpoenaed. It can be breached. The policy that governs it today can change tomorrow. None of that requires bad intent; it is simply the nature of information that lives on someone else’s server.
Where the training data came from.
The models were built on text and images, much of it scraped before consent norms existed. The courts are only now drawing the lines.
In the Bartz v. Anthropic case, a federal court found that training a model on lawfully purchased books can be fair use — while downloading pirated copies to do it is not. The resulting $1.5 billion settlement, granted final approval, is the largest copyright settlement in U.S. history. Both halves matter: the training itself was treated as potentially permissible; the way the data was obtained was not.
The broader point is simpler and harder to litigate. Much of the public web was scraped into training sets before anyone was asked, and as an individual you generally cannot audit whether your words or images are inside a given model. The lawsuits above are the visible edge of a provenance problem that runs underneath the whole technology.
Data brokers meet AI.
AI didn't create the data-broker economy. It lowered the cost of aggregating, inferring, and profiling from it.
“Anonymized” rarely means anonymous. A widely cited Nature Communications study estimated that 99.98% of Americans could be correctly re-identified in any dataset using just 15 demographic attributes. Once data can be re-linked, the distinction between anonymous and identified is mostly theoretical.
The enforcement landscape is starting to respond. The FTC’s orders against the location-data brokers X-Mode and InMarket (2024) and its 2026 settlement with Kochava — banning the sale of sensitive location data — show regulators treating precise location and other sensitive categories as off-limits to sell. AI’s role is the multiplier: it makes aggregating and profiling from broker data cheaper and faster, which is why the same personal details are worth more assembled than apart.
How scattered personal details get assembled into a profile — and weaponized — is the subject of a companion guide.
Doxxing, Disclosure, and AI →Health, genes, and children.
The categories where a privacy failure is hardest to undo are exactly the ones with the thinnest protections.
Genetic data. When 23andMe went bankrupt in 2025, its database of genetic profiles became a transferable asset of the estate. A court approved its sale — about $305 million — to a founder-led nonprofit, over objections from multiple state attorneys general seeking explicit customer consent. The lesson is durable: a privacy promise is only as lasting as the company that made it, and companies do not last forever.
Children’s data.The FTC’s amended COPPA Rule (effective June 2025) now requires separate verifiable parental consent before a child’s personal information is disclosed to third parties — explicitly including for AI training. It is one of the few places the law moved first.
Health and student data.Most people assume HIPAA covers anything health-related. It doesn’t — it covers providers and insurers, not wellness apps or chatbots. FERPA covers school records but has documented gaps for ed-tech. The mental-health chatbot you confide in and the learning app your child uses often sit in those gaps.
A patchwork, not a floor.
There is no comprehensive federal privacy law. What you get depends heavily on which state you live in. A factual map, not a scorecard.
The United States has no comprehensive federal consumer-privacy statute. Instead, sector-specific rules — HIPAA for certain health data, FERPA for education records, COPPA for children, GLBA for financial data — each cover a slice, leaving most everyday AI use uncovered. Into that gap, states have moved: about 20 states have comprehensive consumer-privacy laws in effect as of 2026. Whether you have enforceable data rights can depend on your zip code. As of this writing, North Carolina is not among the states with such a law.
A newer layer of AI-specific state laws is arriving alongside them — Colorado’s AI Act (with a delayed effective date), Texas’s TRAIGA, Illinois’s restriction on AI-delivered therapy, California’s SB 243 for companion chatbots. The picture changes often enough that the honest move is to point to a maintained tracker rather than freeze a snapshot. This section is a map of where the lines are — not an argument about where they should be.
Action for every level of influence.
For yourself
- Check the training setting on every AI tool you use. Consumer versions often use your conversations to improve the model unless you turn it off. It is usually one toggle.
- Use temporary or incognito chat modes for anything sensitive, and delete old conversations where the option exists.
- Assume anything you type may persist — through retention, breach, legal discovery, or a change in the company's policy. Type accordingly.
For families
- Never paste other people's personal information — health details, a friend's address, a child's data — into a consumer AI tool. It is not yours to hand over.
- Review the consent flow on any AI app a child uses. New federal rules require separate parental consent before a child's data goes to third parties, including for AI training.
- Talk about it plainly: the chatbot is not a diary. It is a company's server.
For working professionals
- Use work-provided enterprise accounts for work data. Business and enterprise tiers generally state they do not train on your content by default — consumer tiers often do the opposite.
- Do not solve the data-privacy problem by using an unapproved personal tool at work. That is how sensitive data ends up in the least-protected place.
- If your organization has no AI data policy, ask for one before you need it.
For organizations & schools
- Give people sanctioned tools with clear data terms. Where guidance is absent, people use shadow tools and put sensitive data into them anyway.
- Map which laws actually cover your data. HIPAA and FERPA cover narrow slices; most consumer AI use falls outside them.
- Write down what may and may not be entered into which tools — and keep it current, because vendor practices change.
Related
Algorithmic Bias
Algorithms are making decisions about bail, housing, credit, and healthcare — and their mistakes fall hardest on communities least able to fight back.
AI Scams & Fraud
Scammers are cloning voices, faking emergencies, and building relationships that don't exist. Here's what that sounds like — and how to protect yourself.
AI & Healthcare
Insurers are using algorithms to deny care at scale. What the documented cases show, and how to appeal a decision an algorithm made about you.
Where this leads
CPAI teaches this in workshops and cohort programs.
We deliver this material to schools, libraries, employers, and community organizations — in person and online.
Research & further reading.
Want CPAI to teach AI data safety in your school or workplace?
We deliver this material as workshops and cohort programs for schools, libraries, employers, and community organizations.