Annotated dossier · Compiled September 3, 2026

Evidence & sources

Every claim we make traces to an entry here. We cite the strongest available form of each claim and mark its weight — and we include the studies that cut against us, because a brief that only lists supporting evidence is a brochure.

causalrandomized or quasi-experimental   metameta-analysis or systematic review
surveyrepresentative survey data   reportinstitutional reporting or investigation

The single most important study

The guardrail was the entire difference.

Bastani et al. (2025) gave Turkish high-school students the same underlying model in two configurations, then tested them without it. One arm collapsed. The other matched the textbook control.

Unassisted exam performance under three study conditions Indexed to the textbook control at 100. Students given unrestricted general-purpose AI scored 83 — a 17 percent drop. Students given a tutoring-configured AI that supplied hints rather than answers scored 100, matching the control. 0 25 50 75 100 Textbook control — unassisted exam performance: 100 (baseline) 100 Unrestricted general-purpose AI — unassisted exam performance: 83, a 17% drop 83 −17% Tutoring-configured AI, hints not answers — unassisted exam performance: 100, matching control 100 CONTROL BASELINE Textbook control Unrestricted AI general-purpose Tutoring-configured AI hints, not answers Unassisted exam performance, indexed to control = 100 · Bastani et al. (2025)
View as table
ConditionUnassisted exam performance (indexed)
Textbook control100
Unrestricted general-purpose AI83 (−17%)
Tutoring-configured AI (hints, not answers)100 (parity with control)

Same students. Same subject. Same underlying model. The only difference between the second and third bar is whether the tool was configured to give answers or to give hints. This is why we argue the policy variable is configuration and pedagogy, not proximity — and why a category ban optimises the wrong thing.

AI A researcher at a desk stacked with printed academic papers marked up in pen and flagged with coloured tabs, holding one up to compare against another.
Sixty sources sit behind this site. Of the 800-plus papers in Stanford’s review of AI in K-12, about twenty produce strong causal evidence — and none of those studied students in U.S. schools. Reading the pile is how you find that out.

Section B

Actual child exposure and harm

survey

Pew Research Center — How Teens Use and View AI

Published February 24, 2026 · fielded Sept–Oct 2025 · n = 1,458 U.S. teens 13–17 · ±3.3 pts · Ipsos KnowledgePanel

  • 64% use AI chatbots; ~30% daily
  • 54% have used one for schoolwork; ~1 in 10 for all or most schoolwork
  • 59% think AI cheating happens at least somewhat often at their school
  • 12% have sought emotional support from a chatbot

Cuts both ways: 34% of pessimistic teens cite overreliance and lost critical thinking as their main concern. Teens are worried about exactly what the ban is worried about. We read that as evidence for teaching.

pewresearch.org →

survey

Common Sense Media — Talk, Trust, and Trade-offs

Robb & Mann, 2025 · n = 1,060 teens 13–17, nationally representative

  • 72% have used an AI companion; 52% at least a few times a month
  • 1 in 3 use them for social and relational purposes
  • Younger teens (13–14) trust companion advice more than older teens
  • Only 37% of parents know their teen uses AI

We agree with this source's conclusion. Its risk assessment found companion platforms pose unacceptable risks to under-18 users and recommends none use them until safeguards improve. We support the companion prohibition without qualification.

commonsensemedia.org →

survey

Common Sense Media Census — AI Use by Tweens and Teens (2026)

  • 7 in 10 teens have used a generative AI tool; 40% for school assignments
  • 46% of those did so without teacher permission
  • Policy awareness: 35% say their school has guidelines, 27% say no rules, 37% are unsure

Policy communication is already failing before enforcement begins.

Full census PDF →

report

UNICEF / ECPAT / INTERPOL & the WIRED–Indicator investigation

  • 1.2 million children aged 12–17 across 11 countries reported images manipulated into sexual deepfakes in a single year
  • 600+ victims across ~90 schools in 28 countries since 2023; 94% girls
  • Pennsylvania: two 14-year-olds generated ~350 fake nude images of 59 classmates from school-published photographs
  • July 2026 — UK IWF and Australian eSafety warned schools their public photo galleries are being harvested

Load-bearing for the relocation argument. Perpetrated by minors, on personal devices, outside school networks. A district device ban has no contact with this harm vector.

Section C

The track record of prohibition

meta

CDC Community Guide meta-analysis (2012)

66 comprehensive risk-reduction programs vs. 23 abstinence-only programs. Comprehensive programs improved sexual activity, frequency, contraception use and number of partners. Abstinence-only: insufficient evidence of any change.

Supported by Columbia Mailman analysis and Stanger-Hall & Hall (J. Adolescent Health) finding state abstinence-only policy associated with higher teen pregnancy.

Contrary evidence we include: Jeynes (2020) reports more favourable attitudinal findings for abstinence programs. The behavioural-outcome literature remains dominated by null findings.

report

PEN America — school book ban index

  • ~23,000 bans since 2021
  • 6,870 in 2024–25 across 23 states, 87 districts; 3,743 unique titles
  • 79% of banned titles were written for children and young adults
  • Spring 2026: doubling of nonfiction censorship

Paired with Dennis (Univ. of Rhode Island, 2025) on reading volume as a predictor of achievement. We label that piece a synthesis, not primary research.

pen.org/book-bans →

report

CIPA filtering — two decades of results

Over-blocking of legitimate sexual-health and LGBTQ+ educational content, alongside documented circumvention by adolescents via VPN, alternate DNS and proxy.

Weight note: this literature is largely vendor and practitioner reporting rather than peer-reviewed evaluation. We rely on it for pattern, not for effect size.

causal Counter-evidence

Cellphone bans — the strongest case against us

Largest U.S. study (Univ. of Michigan Youth Policy Lab, 2026): bell-to-bell bans cut in-class phone use from 61% to 13%; well-being improved with duration; but average test-score effects near zero, with little effect on attendance, attention or perceived bullying.

Figlio et al. / NBER (2025) on Florida's statewide ban found test-score gains after an adjustment year, especially for low achievers and males, with significant reductions in unexcused absences.

Any honest reading holds both. Bans can work, inconsistently, on outcomes adjacent to the ones claimed. We do not assert prohibition is uniformly ineffective.

Section D

Inoculation, prebunking and media literacy

The mechanism evidence. This is what the organisation is named after.

causal

Roozenbeek, van der Linden & Nygren — Bad News across cultures

Significant reduction in perceived reliability of tweets using deception techniques, replicating across the U.S., Sweden, Germany, Greece and Poland. Effect size d = 0.37. Confers resistance to strategies, not merely to trained examples.

HKS Misinformation Review · read →

causal

van der Linden et al. — Science Advances (2022)

Psychological inoculation improves resilience against misinformation on social media in large field experiments. Demonstrates the effect survives platform scale.

science.org →

meta Most load-bearing

Media Literacy Education and Misinformation Among Adolescents (2026)

Systematic review and meta-analysis, 18 studies, moderate overall effect. The largest effect sizes were for lateral reading and cognitive inoculation, especially when applied in a guided and contextualised manner within the school environment.

This is the single citation that most directly justifies AInoculate's delivery model — guided, in school, technique-focused.

Journalism and Media 7(2) · mdpi.com →

causal

NewsWise randomized controlled trial (2026)

Effective at developing primary-school children's ability to identify misinformation, with a positive relationship between news literacy and civic engagement.

Directly rebuts the "children this young cannot think critically about this" objection.

Cogent Education · tandfonline.com →

meta

JMIR systematic review & meta-analysis (2023); Huang, Jia & Yu (2024)

Inoculation improves credibility assessment, sharing intention and discernment. Media literacy interventions reduce misinformation belief, improve discernment and reduce sharing.

Also: Traberg, Roozenbeek & van der Linden find post-exposure "therapeutic" inoculation retains benefit — relevant because the children we reach have already been exposed.

report

Retention data

One media literacy course produced a 35% improvement in discerning true from false health headlines. 80–90% of trained students reported using the skills one month later, with most retaining cross-checking behaviour at 18 months.

Section E

What the learning-outcomes evidence actually says

This section contains the findings least favourable to AI in classrooms. We present them in full, because our argument depends on reading them correctly rather than on wishing them away.

meta The anchor citation

Fesler, Martinez Claeys, Agnew & Loeb — The Evidence Base on AI in K-12: A 2026 Review

AI Hub for Education, SCALE Initiative, Stanford University, 2026 · drawn from an October 2025 snapshot of 800+ papers

On the state of the evidence

  • Only ~20 papers produce strong causal evidence
  • No high-quality causal studies of AI and students in U.S. K-12 settings
  • Much research is on adults, international, short-duration and short-horizon

On students

  • Immediate gains with access — math, programming, writing
  • Short-term boost, uncertain transfer when assessed unassisted
  • Easier doesn't mean better — relief of cognitive burden can cost deeper thinking
  • Pedagogical design matters — guardrailed tools show more promise

On educators

  • Less time on lesson prep without reducing lesson quality
  • Automated feedback to human tutors improves instructional quality and student outcomes, especially for less experienced tutors

On equity and wellness

  • Largely unexamined in the causal literature
  • No studies of impacts for students with IEPs or 504 plans
  • Tools optimised for English may underserve English Learners
  • AI social companions raise unresolved safety and development questions

Its conclusion: tools designed to foster independent reasoning are more likely to support durable learning. We read that as a mandate for curriculum design and procurement discipline — not for absence. The review itself endorses no one's position, including ours. Full report PDF →

causal

Bastani et al. (2025) — Generative AI Without Guardrails Can Harm Learning

Charted above. Unrestricted access: −17% on unassisted exams. Tutoring-configured access with hints rather than answers: parity with the textbook control.

Students keep bypassing heuristic guardrails — can we build a "science of guardrails" to understand what works in education?

Dr. Hamsa Bastani, University of Pennsylvania

causal preprint

Kosmyna et al. (2025) — Your Brain on ChatGPT

MIT Media Lab. n = 54, four sessions over four months, EEG plus NLP plus human and AI scoring. LLM users showed the weakest neural connectivity and underperformed at neural, linguistic and scoring levels. 83% could not quote an essay written minutes earlier, vs. 11% of search-engine and unaided writers.

Caveats we state wherever we cite it: preprint, small n, single task type, adult participants. MIT's own explainer cautions against overgeneralisation. We teach it as a replication exercise, not as settled science.

arxiv.org →

causal

The preference trap — Kreijkes (2026), Blasco & Charisi (2025)

Note-taking outperformed AI-only use for comprehension and retention, and students preferred the AI chatbot and rated it most helpful despite it producing the worst learning. Separately, students rated a Socratic chatbot less helpful than one that gave direct answers.

Presented against itself: Degen et al. (2025) found pre-service teachers rated a Socratic chatbot as more supportive of critical thinking — the opposite result. Stanford reads the pair as evidence that preference is design- and context-dependent. We present both.

causal

Further findings from Stanford's summary table

  • Chatbot access had no overall learning effect, increased topics covered, harmed understanding, and widened gaps for students with low prior knowledge
  • Early AI access improved homework but had no significant effect on unassisted exams
  • AI-guided tutors matched human tutors and showed superior transfer to novel topics
  • Step-by-step reasoning significantly outperformed solution-giving

Section F

Equity

  • 15.7 million Americans lack access to high-speed broadband (U.S. Census Bureau, 2026); low-income households disproportionately lack reliable internet and devices
  • Districts diverge sharply — some integrating AI, others instructed to avoid it, many lacking staffing, funding or leadership bandwidth to evaluate tools responsibly
  • Policy energy consumed by integrity concerns and bans addresses real risks but can deepen divides, displacing attention from the equity problem

Sources: edCircuit, GovTech, Frontiers in Computer Science (2026), PA TIMES — origin of the "public capability vs. private advantage" framing.

Section G

Frameworks & policy landscape

  • UNESCO AI Competency Frameworks for Students and Teachers (Sept 2024) — human-centered mindset, AI ethics, AI techniques and applications, AI system design. Both insist AI must support rather than replace human intellectual development.
  • ~71 AI-in-education bills across 27 states in the 2026 session (FutureEd)
  • Idaho, Maryland, Oklahoma and Virginia now require state responsible-use guidance
  • Arizona mandates ethical AI instruction; New Mexico offers AI ethics as a CS elective; Hawaii has directed a statewide K–12 AI literacy curriculum
  • As of Feb 2026, 891 of 1,090 AI-literacy papers in the ACM Digital Library were published in 2024–2026 — but no longitudinal studies exist on long-term outcomes

Section H

Where we may be wrong

Stated plainly, because a brief that only lists supporting evidence is a brochure.

  1. There is no causal evidence that AI-literacy curriculum improves AI-specific outcomes in U.S. K-12. Our mechanism evidence comes from adjacent domains — misinformation inoculation and media literacy — that we argue transfer. That argument is reasoned, not yet measured.

  2. No longitudinal studies exist of AI ethics interventions on long-term behaviour outside school. Our durability claim rests on media literacy retention data, not AI-specific data.

  3. The phone-ban literature genuinely splits. Prohibition is not uniformly ineffective, and we do not claim it is.

  4. The MIT cognitive-debt study is a small-n preprint on adults doing one task type. We cite it as illustrative, not settled.

  5. CIPA circumvention evidence is practitioner reporting, not peer-reviewed evaluation.

  6. Stanford's review is the strongest available synthesis and it declines to endorse anybody's position, including ours. It says the evidence is thin. We agree that it is thin, and we argue from mechanism and from the track record of prohibition where outcome data does not yet exist. Readers should weigh that accordingly.

Our evaluation commitment. Every dose ships with a pre/post instrument measuring unassisted transfer. Partner districts commit to randomized or staggered rollout where feasible. Instruments, anonymized data, and null results are published — especially null results. We will report if this does not work.

Where we are wrong, we would like to know. Corrections to any citation are welcome and will be acknowledged in the repository history. Get in touch →

Why the dossier is public

Argue with us using our own sources.

Every number on this site traces to something on this page, and every entry carries its weight — causal, meta-analytic, survey, or institutional reporting. Where the evidence cuts against us, it is here too, labelled as such.

That is not modesty. A case for teaching children to check claims cannot itself be un-checkable. If a citation here is wrong, weak or misread, we would rather find out from you than from a school board.