What the evidence shows about AI in classrooms — how students' use of it changes learning, and how prepared the adults in the room are to teach that difference.
Education & AI literacy
The same AI tool produces opposite outcomes depending on how it is used — and usage guidance is largely absent while adoption is near-universal.
What’s happening
AI assistants are now in daily use across work, school, and home. The evidence increasingly shows that the outcome — whether AI use builds capability or erodes it — depends less on how much people use it than on how.
What the evidence shows
In a randomized trial of about 1,000 students, unrestricted GPT-4 raised practice performance 48% but left those students 17% worse on a later unassisted exam; a version redesigned to give hints instead of answers eliminated the harm. A separate set of randomized studies found that people who used AI mainly for direct answers underperformed afterward, while those who used it for hints showed no deficit. One positive result, one negative — both randomized. The brief's credibility rests on presenting both.
Well-designed AI tutoring can roughly double learning gains over conventional instruction. The common thread across the evidence is design and usage mode, not access alone.
Where it reaches constituents
Students, workers, and families are adopting AI universally while usage guidance lags. Roughly three in four knowledge workers already use AI at work, and most received no training — so the habits that determine whether AI helps or harms are being formed by default, not design.
The current legal & regulatory landscape
The EU AI Act's Article 4 AI-literacy obligation has been in force since February 2025 — the first legal AI-literacy mandate. There is no comparable federal AI-literacy requirement in the United States. North Carolina's AI Strategic Roadmap commits to foundational AI-literacy training across all 100 counties by 2028.
Considerations policymakers are weighing
- ·How AI-literacy programs define "literacy" — tool mechanics versus the usage habits and evaluation skills that the evidence links to outcomes.
- ·Whether learning products disclose whether they are designed to give hints or answers.
- ·How programs measure outcomes (durable capability) rather than attendance.
Full evidence & citations: cp-ai.org/policymakers/briefs/safe-ai-use
Educators & professional learning
Teachers are being asked to supervise a technology whose effect on learning depends entirely on how students use it — and the guidance they receive is thinnest for exactly that job.
What’s happening
Student use of AI for schoolwork has risen sharply in three years while formal guidance for the teachers supervising it has moved far less. The guidance that does exist concentrates on the teacher's own productivity and on detecting AI-written work. Almost none of it addresses the job the evidence says decides the outcome: teaching students to use AI in ways that build capability instead of replacing it.
What the evidence shows
A nationally representative survey of 2,069 public K-12 teachers fielded in February and March 2026 found 82% receive no formal guidance on applying AI tools to their work. The pattern inside that number is the point: “no guidance” rates were highest for the student-facing uses — 69% for one-on-one instruction and tutoring, 71% for coaching their own practice — and lowest for producing worksheets and assignments, at 47%. Guidance is most available where AI saves a teacher time, and least available where it touches a student. (Survey commissioned by a foundation that promotes AI adoption in K-12; fieldwork was conducted on a probability-based national teacher panel.)
Separate national polling of 806 sixth- through twelfth-grade teachers describes what the training that does occur actually covers. Across all of them, only 11% say “how to respond if you suspect a student's AI use is detrimental to their well-being” was covered in any AI training or information their school provided — the least common topic asked about, below detecting AI-generated student work (18%) and responding to suspected misuse (15%). Teachers' own top-ranked priority was guidance on using AI tools effectively, and that was the most commonly covered topic at 29%. On the student side, more than 80% of students report their teachers did not explicitly teach them how to use AI, and 30% of teens say a teacher has ever discussed how to use AI safely.
This matters because the instructional evidence identifies teacher design as the variable that decides the outcome. In a randomized trial of roughly 1,000 high school students, unrestricted GPT-4 raised in-session performance 48% but left those students 17% worse on a later unassisted exam; a version whose safeguards were written by two of the school's own math teachers eliminated that harm entirely. And in a two-year randomized trial across 18 middle schools, an AI tutor deliberately configured to coach rather than give answers produced gains resembling the same platform without AI — because 96% of students tried it, but the median student messaged it on only a third of the days they practiced, and in only 17% of the sessions where they made a mistake. The tool was built correctly. Nothing taught the students, or the adults around them, to use it that way.
Where it reaches constituents
Teachers in every subject and grade, and the students they supervise. Because the effect of student AI use turns on usage mode rather than access, the adult best positioned to change the outcome is the classroom teacher — and that is the role current guidance addresses least. The gap is also uneven: district-level AI training reached 67% of low-poverty districts in fall 2024 against 39% of high-poverty districts, and nearly all of it was optional.
The current legal & regulatory landscape
Federal Executive Order 14277 (April 2025) directs the Secretary of Education to prioritize existing discretionary grant funds for educator AI training and to support professional development for educators across subject areas. It creates no new appropriation and requires nothing of states or districts. At least 27 states had issued K-12 AI guidance as of August 2025; that guidance is advisory in every case reviewed, and several state documents say so explicitly.
A much smaller number of states have legislated anything about educator AI training, and one has gone furthest on what that training must look like. Maryland's Artificial Intelligence Ready Schools Act, signed May 26 and effective June 1, 2026, requires the state education department to provide professional development to educators and school leaders that is “sufficient in duration and quality” and substantively covers effective use, privacy, security, academic integrity, and means of avoiding overdependence — delivered either during the regular workday as designated professional-development time or as a course eligible for licensure-renewal credit. The duty runs to the department to offer it, not to the individual teacher to complete it. The same Act requires every local school system to adopt an AI policy within 120 days of the department's guidance and to designate an AI coordinator. Idaho and Hawaii have enacted measures including training provisions, Hawaii's with a $5 million pilot. No rigorous evaluation of an AI-specific educator training program with measured student outcomes appears to exist: a Stanford review screened more than 800 papers and identified 20 high-quality causal studies, none of them evaluating AI teacher professional development.
Considerations policymakers are weighing
- ·Whether educator AI training addresses the instructional mediation of student AI use, or is bounded by tool mechanics and academic-integrity policing — the two areas the national data show currently dominate.
- ·How the format of any training requirement compares with what the professional-development research associates with measurable effect. A review of 35 rigorous studies identifies content-specific focus, coaching, collaboration, modeled practice, and sustained duration; NCDPI's own guidance recommends a four-to-six-week practice period. Short asynchronous module formats are the most common approach and the least studied.
- ·Whether a training requirement is paired with evaluation, given that no rigorous evaluation of AI-specific educator professional development currently exists to copy.
- ·How curriculum requirements for students are sequenced relative to preparation for the teachers who deliver them.
Full evidence & citations: cp-ai.org/policymakers/briefs/educator-ai-readiness
Learning & cognition
AI can raise a student's performance while lowering their learning — and the gap between the two is invisible from inside the classroom.
What’s happening
Students use AI for schoolwork at scale, and the measurable question is no longer whether they use it. It is what happens to learning when they do. The research converges on an uncomfortable finding: performance during the task and capability after it can move in opposite directions, and the difference is decided by whether AI removes the effort or supports it.
What the evidence shows
The clearest result comes from a randomized trial of roughly 1,000 students across four practice sessions. Unrestricted GPT-4 access raised in-session grades 48%. On the subsequent unassisted exam, those same students scored 17% lower than classmates who never had access. A second version of the same model — instructed to give hints rather than solutions, and loaded with common student mistakes and the matching feedback — raised in-session grades 127% and produced essentially no effect on the unassisted exam. The harm was removed by pedagogical design, not by a better model.
Two 2026 field experiments sharpen this into a practical finding about structure. A two-year randomized trial across 18 middle schools, using an AI tutor explicitly configured to coach rather than answer, raised math achievement roughly 0.06 to 0.08 standard deviations over a school year — gains the researchers describe as resembling the same practice platform without AI at all. Their explanation is engagement: students rarely used it as a tutor. A companion randomized experiment with more than 6,000 students found AI's clearest benefit appeared specifically after mistakes, improving next-attempt correctness and shortening the path back to a correct answer. Delayed-test gains appeared only where AI was embedded in a mastery structure that forced engagement at the point of error.
The reason this is hard to manage locally is that no one in the room can feel it. A preregistered experiment found that any AI involvement impaired participants' memory a week later of which ideas had been their own. Preregistered studies with 2,691 participants found people systematically underestimate how much they rely on AI and overestimate what it saves them. Learning science supplies the mechanism: the practice conditions that feel most fluent reliably produce the least durable learning, and AI is very good at removing exactly the friction the learning was made of. Students themselves register the concern more than administrators do — 67% of surveyed youth agreed that more student AI use will harm critical thinking, against 22% of district leaders.
Where it reaches constituents
Students at every level, and the teachers and families who cannot see the trade-off from outside it. The cost shows up only once the tool is taken away, in what the student can still do without it. That is also why individual course-correction is unreliable here: the signal people would need in order to adjust is precisely the one the research says they do not receive.
The current legal & regulatory landscape
This is largely unregulated territory. Federal Executive Order 14277 (April 2025) directs existing discretionary funds toward AI education without new appropriation. State activity concentrates on advisory guidance for districts and on student-facing curriculum requirements. Standards for how AI learning products are designed or evaluated are almost entirely absent.
No federal or state requirement currently exists that a learning product disclose whether it is built to give hints or answers — the single design variable the randomized evidence identifies as decisive. The evidence base itself is also thin: a Stanford review screened more than 800 papers and found only 20 high-quality causal studies, most of them short-term and measured on practiced material.
Considerations policymakers are weighing
- ·Whether AI learning products disclose the usage mode they are designed for — hints versus answers — since that is the variable the randomized evidence identifies as decisive.
- ·Which outcome measures count in evaluating a learning tool: performance during use, or performance once the tool is removed.
- ·Whether existing health-and-wellbeing instruction requirements are a natural home for AI usage content, or whether it belongs in computer science.
- ·Support for longitudinal measurement, which barely exists — nearly all current studies are single-semester and measured on practiced material.
Full evidence & citations: cp-ai.org/policymakers/briefs/ai-and-student-learning