
Share
A Turkish classroom experiment offers a cautionary data point for schools embracing chatbots: students who leaned on ChatGPT solved more practice problems but scored significantly worse once the AI was taken away.
Picture a student riding a bike with training wheels for months, never once removing them. They feel steady, confident, fast even. Then the training wheels come off for the big race, and suddenly they can barely stay upright. That, in essence, is what researchers at the University of Pennsylvania found when they studied nearly 1,000 high schoolers in Turkey using ChatGPT to help with math practice.
The findings, first reported by The Hechinger Report, deserve attention from anyone who cares about how young people learn, not just how they perform in the moment. Because the gap between those two things turned out to be enormous.
Here is what the researchers did. They split students into three groups while working through math practice problems. One group had access to ChatGPT. Another used a specialized AI tutor built on ChatGPT, designed to guide rather than simply answer. A third group practiced the old-fashioned way, with no AI assistance at all. Then everyone took a test without any AI tools in hand.
During practice, the AI groups looked like they were thriving. Students using plain ChatGPT solved 48% more problems correctly than their unassisted peers. Those using the AI tutor did even better, solving 127% more problems correctly. On paper, that looks like a stunning educational win. Give a kid a smart assistant, watch their accuracy soar.
But the test told a different story. Students who had used ChatGPT during practice scored 17% worse on the follow-up test than students who had practiced without any AI help. The tutor group did not fare much better once the training wheels came off. Meanwhile, students who practiced entirely on their own performed almost identically on their practice work and their tests. No dramatic rise, no dramatic fall. Just steady, consistent performance that reflected what they actually understood.
It is worth pausing on why this happened, because the mechanism matters as much as the result. Researchers told The Hechinger Report that students were using the chatbot as a crutch, and that this reliance can substantially inhibit learning. That word, crutch, is doing a lot of work here. A crutch helps you walk when you are injured. It does not teach your leg to heal faster. In fact, lean on it too long and the muscles you need for independent walking can atrophy from disuse.
Something similar seems to be happening with math skills. When ChatGPT hands a student a correct answer or walks them through a solution step by step, the student experiences the satisfaction of getting it right. But that satisfaction can be misleading. Getting a problem right because a tool solved it for you is not the same as building the mental scaffolding to solve the next problem alone. The practice session created an illusion of mastery that evaporated the moment the AI was removed.

This distinction matters enormously for how we think about technology in classrooms. Educators have spent years trying to figure out how calculators, spreadsheets, and search engines fit into learning without undermining it. AI chatbots raise the stakes considerably, because unlike a calculator, which performs a narrow, well-understood function, a chatbot can essentially do the entire intellectual task a student is supposed to be doing themselves. The practice problems were meant to build durable understanding. Instead, for a meaningful share of students, they became an exercise in outsourcing.
None of this means AI tutoring tools are useless. The 127% improvement during practice shows these tools genuinely can walk a student through difficult material and help them produce correct work. That is not nothing. For a student stuck and frustrated, immediate, patient guidance has real value. The danger lies in mistaking that guided success for independent competence, and in designing homework or practice routines that never ask students to demonstrate they can do the work without assistance.
There is also a broader question lurking here about how we measure learning in an AI-saturated world. If practice scores look great but test scores collapse, which number should a school trust? Which one should a parent trust when checking in on their child's progress? This experiment suggests that practice performance alone, when AI is involved, may no longer be a reliable signal of what a student has actually absorbed. Teachers and parents assessing progress through homework completion rates could be seeing a rosier picture than reality warrants.
This single Turkish high school experiment will not settle the debate over AI in education, and it should not be treated as the final word. But it adds a concrete, measurable data point to a conversation that has, until now, leaned heavily on intuition and anecdote. Nearly 1,000 students, a controlled comparison across three groups, and a clear, consistent pattern: more AI help during practice correlated with worse performance when that help disappeared.
For school districts weighing how aggressively to integrate chatbots into daily instruction, this is a flashing caution light rather than a stop sign. The tools can help. They can also quietly erode the very skills they are meant to support, especially when students use them to shortcut the struggle that learning actually requires. Struggle, frustrating as it feels in the moment, appears to be doing real cognitive work that a quick AI-generated answer simply cannot replace.
As more classrooms adopt these tools, the challenge will be designing practice and assessment that preserves the benefits of AI assistance without letting it become a substitute for genuine understanding. Getting that balance wrong risks producing a generation of students who look capable on paper but struggle the moment real-world problem solving demands they work it out themselves.
Tags
Original Sources
UPDATE: Students using artificial intelligence did worse on tests, experiment shows
↗ https://edsource.org/updates/students-using-artificial-intelligence-did-worse-on-tests-experiment-shows
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
2 October 2026
28 articles
Related Articles

County Officials Get a Crash Course in AI, But Questions Remain About Who Gets Left Behind
Job Market & Society · 5 min

The All of Us Data Gap Isn't About Trust. It's About Turning Records Into Usable Data
Job Market & Society · 6 min

Stanford HAI Names 15 PhD Researchers to New Data Science Scholars Cohort
Job Market & Society · 5 min
Related Articles

County Officials Get a Crash Course in AI, But Questions Remain About Who Gets Left Behind
Job Market & Society · 5 min

The All of Us Data Gap Isn't About Trust. It's About Turning Records Into Usable Data
Job Market & Society · 6 min

Stanford HAI Names 15 PhD Researchers to New Data Science Scholars Cohort
Job Market & Society · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.