Skip to main content

Study guide

AI study tools: help or harm? What a large classroom trial found

A large trial found unrestricted AI tutoring lifted practice scores but lowered exam scores, while guardrails removed the harm. What that means for how you study with AI.

Published 2026-08-09 · 6 min read · Captivate Exam Hub

The question isn't whether you'll use AI. It's how.

Most students preparing for senior exams already use an AI chatbot somewhere in their study: to explain a concept, check working, or summarise notes. The interesting question is no longer whether that's allowed or common. It's whether it helps you on exam day, and the best evidence so far says: it depends heavily on how the AI behaves.

What the trial found

A large randomised classroom study published in PNAS put this to the test with high school maths students [7]. Students were split into groups: some practised with a standard AI chatbot that would provide answers on request, some with a tutor version constrained to give hints and guidance without revealing answers, and some with no AI at all.

The headline results:

  • Both AI groups did substantially better than the no-AI group during practice. The help genuinely helped, while it was there.
  • When the AI was taken away for the exam, students who had used the unrestricted chatbot performed worse than students who had never had AI at all. The practice gains reversed into an exam loss.
  • Students who used the guardrailed tutor showed no such harm. The safeguards, not the AI itself, made the difference.

One more finding is worth sitting with: students with the unrestricted chatbot often believed they were learning well. The tool was functioning as a crutch while feeling like a teacher.

Why answers-on-tap backfire

The result stops being surprising when you connect it to the retrieval-practice literature. Learning happens when you generate an answer: the struggle to produce it from memory is what strengthens it [3]. An AI that hands over a polished solution the moment you're stuck removes exactly that step. You read the solution, it makes sense, it feels like progress. But reading a correct answer is re-reading with extra steps, and re-reading is the technique that keeps losing in controlled studies.

There's a second, sneakier cost: an always-available solver hides your gaps from you. If every stall gets smoothed over instantly, you never find out which topics you can't actually do unaided, and neither does your revision plan.

What good AI help looks like

The same trial points at the fix. AI help that preserves the generation step, making you attempt first and nudging rather than answering, kept the practice benefit without the exam harm [7]. In practice that looks like:

  • An attempt gate: no help until you've had a genuine go at the question.
  • Escalating hints: strategy first (“what kind of problem is this?”), then a nudge at the next step, never the full answer while the question is live.
  • Socratic questions back at you, so you stay the one doing the reasoning.
  • Explanations after you've answered, when they consolidate rather than replace.

Using a general chatbot without hurting yourself

You don't need a special tool to apply this. If you study with a general chatbot, a few self-imposed rules recover most of the safeguards:

  • Attempt every question before you ask. Paste your working, not just the problem.
  • Ask for a hint, not a solution: “don't give me the answer — tell me what to look at next.”
  • Ask it to quiz you on a topic, then answer from memory before checking.
  • Verify anything that matters against your official study design and course materials. AI-generated content can be wrong, confidently.

How we apply this (and our own caveat)

This study is the reason Captivate Exam Hub's AI hints work the way they do: solutions stay locked until you've attempted, and the hint ladder climbs one rung at a time, strategy before nudge, and never hands over the answer. The goal is the guardrailed condition from the trial, not the unrestricted one.

The same honesty applies to us as to any AI tool: AI-generated questions, feedback and summaries can contain errors, so check important points against your official study design. Our AI Use & Accuracy notice spells that out. And if you want the broader evidence base these features sit on, start with our guide to evidence-based revision.

References

  1. [3] Karpicke (2025), review of retrieval-based learning — practice testing outperforms restudy across hundreds of experiments.
  2. [7] Bastani et al. (2025), PNAS — unguardrailed AI tutoring improved practice but lowered exam performance; guardrails removed the harm.