
Share
A Turkish high school experiment found students using ChatGPT solved far more practice problems correctly, then scored worse on tests than peers, raising hard questions about what AI tools actually teach.
Picture two students studying for the same math test. One works through practice problems alone, struggling, erasing, trying again. The other has ChatGPT open in a browser tab, ready to supply an answer the moment things get hard. Which one do you think understands the material better when the test arrives?
According to new research covered by The Hechinger Report, the answer is not the one you might expect if you think of AI as a simple shortcut to better grades. In fact, it is the opposite.
Researchers from the University of Pennsylvania ran an experiment with nearly 1,000 high schoolers in Turkey, and what they found should give parents, teachers and school administrators real pause. Some students practiced math problems with access to ChatGPT. Others used a specially designed AI tutor built on ChatGPT. A third group worked through the same problems entirely on their own, without any AI assistance.
The differences during practice were dramatic. Students using plain ChatGPT solved 48% more problems correctly than those working solo. Students using the AI tutor did even better, solving 127% more problems correctly. On paper, that looks like a resounding win for AI-assisted learning.
Then came the actual tests, given without any AI support. The students who had relied on ChatGPT during practice scored 17% worse than their unassisted peers. The students who worked alone during practice performed just as well on the test as they had during practice, showing no drop-off at all.
Think of it like learning to walk with a set of training wheels that never come off. You feel steady, you move quickly, you avoid falling. But the moment those wheels are removed, you discover your balance was never really yours to begin with.
That is essentially the metaphor researchers used when describing what happened to the ChatGPT users in this study. They told The Hechinger Report that students treated the chatbot as a "crutch," something to lean on rather than a tool to build strength with. And crutches, however useful in the moment, do not build muscle. They can, in the researchers' words, "substantially inhibit learning."

This distinction matters enormously for how we think about AI in classrooms. There is a meaningful difference between a tool that helps you practice a skill and a tool that does the skill for you while you watch. When ChatGPT supplied answers during practice sessions, students appear to have absorbed very little of the underlying reasoning. They got the right answer on the page, but not the right process in their heads.
The AI tutor group presents an even more puzzling case. That tool was specifically designed to guide students through problems rather than just handing over solutions, and it produced the largest practice gains of any group in the study. Yet those students still stumbled on unassisted tests, though the report does not specify their exact test scores relative to the ChatGPT-only group. The fact that even a more pedagogically careful AI tool failed to translate into durable, independent skill suggests the problem runs deeper than just "bad" chatbot design. It may be about how any external answer-generating system interacts with the slow, effortful process of actually learning math.
This should resonate with anyone who has ever used GPS navigation to get somewhere new, only to realize afterward they have no idea how to get there again without pulling out their phone. The route existed, and you followed it successfully. But you never built the mental map yourself. Math practice with an always-available answer key can work the same way.
There is also something worth acknowledging here about how tempting these tools are, especially for students juggling homework, extracurriculars and the pressure to perform well. Getting more problems right in the moment feels like progress. It looks like progress on a homework tracker or a parent's radar. But this study is a pointed reminder that the metrics that make us feel good in the short term, like completion rates or correct answers during practice, are not always the metrics that predict real understanding.
None of this means AI has no place in education. The researchers were careful to note that the tool matters and so does the way it is used. An AI tutor built for guided discovery still produced strong practice results, even if it did not solve the transfer problem entirely. The question schools and policymakers now face is not whether to allow AI tools, but how to design and deploy them so students build skills rather than borrow answers.
For students, parents and educators trying to make sense of AI's rapid arrival in classrooms, this experiment offers a concrete, sobering data point rather than speculation. It is easy to assume that any tool improving short-term performance is helping learning overall. This study shows that assumption can be dangerously wrong, at least for skills like math that require sustained, independent practice to master.
As schools across California and the country grapple with how to integrate AI responsibly, without banning it outright or embracing it uncritically, findings like these deserve serious weight. The stakes are not abstract. They involve whether the next generation of students genuinely learns to reason through problems, or simply learns to ask a chatbot for the answer and hope nobody checks their work later. Getting this balance right will take more than good intentions. It will take careful design, honest evaluation and a willingness to accept that easier is not always better when it comes to how children learn.
Tags
Original Sources
UPDATE: Students using artificial intelligence did worse on tests, experiment shows
↗ https://edsource.org/updates/students-using-artificial-intelligence-did-worse-on-tests-experiment-shows
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
25 September 2026
31 articles
Related Articles

AI Coding Tools Added $942 Million to Hospital Bills Without Sicker Patients, Blue Cross Analysis Finds
Job Market & Society · 6 min

Healthcare's Layoff Wave Rolls On: Thousands of Jobs Cut as Hospitals Chase "Sustainability"
Job Market & Society · 6 min

OpenAI Agent Breach of Australia's Medicare Sparks Push for AI Safety Laws
Policy & Regulation · 5 min
Related Articles

AI Coding Tools Added $942 Million to Hospital Bills Without Sicker Patients, Blue Cross Analysis Finds
Job Market & Society · 6 min

Healthcare's Layoff Wave Rolls On: Thousands of Jobs Cut as Hospitals Chase "Sustainability"
Job Market & Society · 6 min

OpenAI Agent Breach of Australia's Medicare Sparks Push for AI Safety Laws
Policy & Regulation · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.