Center for Practical AI
Healthy AI Use · Guide 1The intent hypothesis

Same tool. Opposite outcomes.

In the space of about a year, randomized studies found both that AI help during practice hurt students on the later unassisted exam, and that an AI tutor roughly doubledlearning gains. Same technology. Same subject areas. The difference wasn’t the model — it was whether the AI gave answers or hints.

16 min read · Includes an interactive: Mode Check

The distinction

Not how much. How.

The intent hypothesis: what determines whether AI use erodes your capability is not exposure to AI but the mode you use it in.

Extractionis asking for the finished thing. The essay, the answer, the working code, the decision. You describe the problem and the output arrives complete. Your contribution is the specification and, if you’re careful, the check at the end.

Scaffolding is asking for help with an attempt you are already making. A hint toward the next step. A critique of your draft. An explanation of the error you just hit. The AI shapes the work; you still produce it.

These look almost identical from the outside. Same tool, same session length, often the same person within the same hour. The difference is invisible in the output and enormous in what it leaves behind in you.

89% → 76%

problems solved unaided after ten minutes of AI-assisted work

Liu et al. 2026, preprint

1% → 8%

problems skipped without any attempt, same study

Liu et al. 2026, preprint

≈2×

learning gains from a tutor designed never to reveal solutions

Kestin et al. 2025

of student AI conversations are direct answer-seeking

Anthropic Education Report

The evidence

What each study can actually claim.

These findings are not equally strong, and stacking them without saying so would be exactly the kind of thing this series exists to argue against.

Randomized · system-level · strongest

The high-school study: unrestricted AI help raised practice scores and lowered exam scores

Roughly a thousand students. Those given unrestricted GPT access during practice performed markedly better while they had it — and markedly worse than the control group on the later unassisted exam. Critically, a second version of the same tutor, restricted to hints rather than answers, eliminated the harm. The intervention was the system's design, assigned at random.

Randomized · system-level · strongest

The physics course: a never-reveal-solutions tutor roughly doubled learning gains

In a randomized comparison against active-learning classroom instruction, students working with an AI tutor built never to hand over the solution learned roughly twice as much, in less time. Same underlying technology as the study above. Opposite result. The design is the variable.

Randomized exposure · self-selected mode

The persistence trials: after ten minutes of AI help, people attempted less and solved less

Three studies, 1,222 participants. AI assistance was assigned at random; the AI was then removed. The assisted group solved fewer problems unaided and — the finding the authors flag as most important — skipped substantially more without attempting them at all. Within that group, people who had used the AI for hints showed no deficit while people who had extracted answers showed both effects. But they chose their own mode, so we cannot separate the mode from the person. Preprint, under review.

Preprint · small sample

Skill formation: developers learning with AI scored lower on later mastery

Developers learning an unfamiliar library with AI assistance scored substantially lower on a subsequent mastery test, with the largest gap in debugging — the part that most requires having built a mental model. Engaged interaction patterns preserved learning. Small sample, not yet peer-reviewed.

Usage data · descriptive

In the wild, roughly half of student AI use is direct answer-seeking

Not an experiment and not a claim about harm — just the base rate. If mode is the variable that matters, it is worth knowing which mode is the default.

How to read this stack

Taken together, these results are consistent with, not proof of, the intent hypothesis. The randomized evidence establishes that a system giving hints produces different outcomes than a system giving answers. It does not establish that a person choosing hints gets the same protection — that step is the leap, and it is currently supported only by data where people picked their own mode. Two of the studies above are preprints. One has a small sample. We are telling you this because a series about honest AI use that overstated its own evidence would be a bit rich.

The mechanism

The generation effect.

This part isn't new, isn't contested, and has nothing to do with AI. It's one of the most replicated findings in the study of learning.

Producing an answer yourself encodes it. Reading the same answer does not — or does so much more weakly. Psychologists have been demonstrating this since the 1970s with word pairs, and it has held up across essentially every kind of material anyone has tried it on. The effort of retrieving or constructing something is not an unfortunate cost of learning. It is the learning.

This is why the hint-versus-answer distinction does so much work. A hint preserves your generation attempt — you still have to produce the step. A complete answer converts the problem into a reading exercise. The task looks the same on your screen. Cognitively, one of them is the thing that builds capability and the other is a demonstration of it happening to someone else.

It also explains the shape of the persistence result. What dropped fastest wasn’t accuracy — it was attempting. As the researchers put it, AI “conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own.” Once you stop generating, there is nothing left for the effect to act on.

The honest part

What nobody has shown yet.

1

No study randomizes intent

Every strong result here randomizes what the system does, not what the user is trying to do. Somebody needs to assign people to extraction and scaffolding conditions and see whether the effect survives. Until then the central claim of this guide is a hypothesis with good support, not a finding.

2

Cumulative real-world erosion is conjecture

The measured effects are immediate — same session, same day. Nobody has followed people for two years to see whether extraction habits compound into durable skill loss. It is plausible. It is not demonstrated.

3

The effects that exist are small to moderate

Thirteen percentage points on an unaided problem set is a real effect and not a catastrophe. Treat this as a reason to choose deliberately, not a reason to panic.

4

Self-selection runs through the user-level evidence

The people who ask for hints may simply be the people who were going to persist anyway. Disposition and mode are tangled together in every dataset we currently have on real users.

This is the hypothesis our field — and CPAI’s own research program — is testing. We would rather tell you that than sell you a certainty we don’t have.

Practice

The scaffolding moves.

Four phrasings that change the mode without changing the tool. Copy them, use them, adapt them.

Critique, don't fix
Here's my attempt: [paste your draft] Don't rewrite it. Tell me what's weakest and why, in order of how much it matters. I'll do the revising.

Keeps you as the producer. You get the diagnosis; the repair is still your work, which is where the learning lives.

Hint, not answer
I'm working on this and I'm stuck: [describe the problem and where you've got to] Give me a hint toward the next step — not the answer, and not the whole path. Just enough to get me moving again.

This is the phrasing that, built into a tutoring system, eliminated the harm in the randomized studies.

Explain the concept, let me retry
I hit this error: [paste the error] Explain the underlying concept I'm evidently missing. Don't give me the corrected version — I want to try again first.

Errors are the highest-value learning moments you get. Handing over the fix spends that moment on nothing.

Argue against me
Here's the conclusion I've reached: [state your position] Before I commit to it, make the strongest case against it. Be specific about what evidence would change my mind.

Also the antidote to sycophancy — models trained on human approval agree far more readily than an honest colleague would.

The stakes rule

Extraction is fine — genuinely fine — for low-stakes output you will never need to produce unaided. Reformat this table. Draft this scheduling email. Summarize this thread. Nobody is building a skill there and nobody needs to. Scaffolding is for anything you are trying to learn, or trying to keep. The mistake isn’t using extraction. The mistake is using it by default, on everything, without ever deciding which kind of task you’re looking at.

Interactive

Which mode do you actually use?

Eight realistic moments. For each one, pick the prompt closest to what you’d genuinely type. Some of them should be extraction — the tool will say so.

Take the Mode Check →
What you can do

Action for every level of influence.

1

For yourself

  • Before you open the chat, decide which this is: output you need, or a skill you're building. Extraction is the right call for the first. It is the wrong call for the second.
  • Run one hints-only day. For anything you're trying to learn, ask for the next step rather than the finished thing, and notice how much longer it takes and how much more of it stays.
  • When you catch yourself pasting in a problem and pasting out an answer, add one sentence: "Here's what I think it is — tell me where I'm wrong."
2

For knowledge workers

  • Separate the deliverable from the capability. Nobody needs you to write meeting notes unaided. Somebody eventually needs you to reason through the thing the notes are about.
  • In review, ask for the attempt as well as the output. A colleague who can show you their draft and the AI's critique of it has learned something the finished document doesn't record.
  • Treat "I'd never be able to do this myself now" as data, not as a joke.
3

For educators

  • Allow hints, prohibit answers — and say so as a rule about mode rather than a rule about tools. This is the intervention with the strongest randomized evidence behind it.
  • Target both misconceptions at once. "Using AI to learn is cheating" and "any AI use is fine" are wrong for the same reason: both treat exposure as the variable.
  • Adults need the workplace framing (deliverable vs. skill); students need the exam framing (the day the tool is gone).
4

For policy

  • Require learning products to disclose hint-versus-answer design. In the randomized studies, that single design choice is the difference between harm and benefit.
  • Fund the study nobody has run: randomize usage intent, not just system design, and follow users longer than a single session.
  • Do not regulate AI in education by volume of use. The evidence does not support quantity as the variable.

Where this leads

Reading is one thing. Practicing it is another.

The Applied AI Certification builds practical AI fluency across all six domains — the working competence that advances toward proficiency, with structured practice, feedback, and a cohort on the same problems.

Sources

Research & further reading.

Randomized controlled trialBastani, Bastani, Sungu, Ge, Kabakcı & Mariman (2025)Generative AI Can Harm LearningPNAS. Randomized trial, ~1,000 students across ~50 Turkish high-school math classes. Unrestricted GPT-4 access raised practice performance 48% but lowered the later unassisted exam 17% versus never-AI controls. A teacher-designed, hints-only tutor version raised practice performance 127% — and eliminated the exam harm.
Randomized controlled trialKestin, Miller, Klales, Milbourne & Ponti (2025)AI Tutoring Outperforms In-Class Active LearningScientific Reports. Crossover randomized trial, ~190 Harvard intro-physics students. A GPT-4 tutor engineered to work one step at a time, never reveal full solutions, and encourage attempts first produced roughly double the learning gains of in-class active learning, in less time. Two lessons, elite sample, carefully engineered tutor — not raw ChatGPT.
Randomized trial · preprint, not yet peer-reviewedLiu, Christian, Dumbalska, Bakker & Dubey (2026)AI Assistance Reduces Persistence and Hurts Independent PerformanceThree randomized studies, N=1,222. After roughly ten minutes of AI-assisted work, people solved fewer problems unaided (89% → 76%) and skipped more without attempting them at all (1% → 8%).
Randomized trial · preprint, not yet peer-reviewedLiu et al. (2026), usage-mode analysisHint-seekers vs. answer-seekers within the persistence trialsWithin the second persistence experiment, 61% of participants said they used AI mainly for direct answers, 27% for hints and clarification. The groups were indistinguishable before AI use — but afterward, answer-seekers underperformed the no-AI control (d=0.36) and skipped more, while hint-seekers showed no deficit at all. Mode was self-chosen, so disposition and mode cannot be separated.
Randomized trial · preprint, not yet peer-reviewedShen & Tamkin (2026)How AI Impacts Skill FormationRandomized trial, 52 mostly-junior Python engineers learning an unfamiliar async library, half with an AI sidebar. The AI group scored 17 points lower on the mastery quiz (50% vs 67%, d=0.74), with the largest gap in debugging — while their ~2-minute speed gain was not significant. Screen recordings showed delegation patterns scoring under 40% and conceptual-inquiry patterns scoring 65% or higher. Small sample, immediate assessment, preprint.
Foundational researchSlamecka & Graf (1978)The Generation Effect: Delineation of a PhenomenonJournal of Experimental Psychology: Human Learning & Memory, 4(6), 592–604. Five experiments: self-generated answers were remembered better than read ones, across recognition, recall, and encoding rules. The mechanism predates AI by decades — a complete AI answer turns every problem into a read trial; a hint preserves the generation attempt.
Industry usage reportAnthropic (2025)Anthropic Education Report: How University Students Use ClaudePrivacy-preserving analysis of ~574,000 anonymized student conversations. Nearly half of interactions were direct answer- or output-seeking rather than collaborative — and students most often delegated the highest-order skills, Creating (39.8%) and Analyzing (30.2%). Observational, one product, vendor self-analysis.
Last reviewed: July 2026We review this page quarterly. Statistics in this category change rapidly.The persistence trials and the skill-formation study are preprints and are labeled as such throughout. The user-level hint-versus-answer split is self-selected, not randomized. No study yet randomizes usage intent.

Want CPAI to teach this in your school or workplace?

We deliver the Healthy AI Use material as workshops and cohort programs for schools, libraries, employers, and community organizations.