
Share
As artificial intelligence models become more sophisticated, they still lag behind human children in learning language efficiently. This gap holds critical insights for both AI and education.
In the realm of language, a fundamental aspect of human interaction, there’s an intriguing new player that has joined the ranks: artificial intelligence. For centuries, only human children could learn languages to perfect fluency. Now, large language models (LLMs) like Claude, DeepSeek, and OpenAI’s GPT series can converse naturally with us. Yet, despite their sophistication, these machines still fall short in one crucial area: data efficiency.
Four years after ChatGPT's debut, the ability to chat with our devices has become almost second nature. These LLMs are so advanced that they can mimic human conversation convincingly. However, beneath this facade lies a significant issue: these models require an astronomical amount of data to achieve fluency. An LLM might process a hundred thousand times more words than a person encounters while mastering their native language-a stark contrast to the relatively modest exposure children have by their first birthday.
Michael C. Frank, a cognitive scientist at Stanford University, acknowledges the impressive progress in AI but points out the inefficiency: “The progress recently has been amazing, but we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.”
This disparity is known as the data efficiency gap. It highlights how children can outperform even the most advanced AI models with far less input. A preteen raised in a linguistically rich environment might hear around 100 million words, and by age 20, this number could reach up to 300 million words if literacy is factored in. In contrast, modern LLMs like Meta’s Llama 3.1 have been trained on an estimated 15 trillion tokens (word-like chunks of language) during pretraining.
Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University, explains the scale: “Claude has seen the amount of language that an entire city will experience in one generation.” If you were to print out all the words used to train a modern LLM, the stack would be taller than most buildings.

The data efficiency gap raises important questions for both AI research and cognitive science. For the past decade, improvements in language models have largely come from increasing their size and training them on more data. However, this approach is not sustainable indefinitely. Meta’s Llama 3.1 already required a massive amount of data, and future models could require ten times more. Yet, there is only so much internet content available for training, and the well could run dry as early as the 2030s.
Understanding how children learn language with such efficiency could unlock new avenues for AI development. Cognitive scientists are exploring whether there are underlying principles or mechanisms that allow human brains to process language more effectively. If these can be identified and replicated, it could lead to more data-efficient AI models that require less computational power and fewer resources.
The insights gained from studying children’s language acquisition could also benefit education. In a world where technology is increasingly integrated into learning environments, understanding how humans naturally learn languages could help design better educational tools and curricula. This could be particularly important in addressing the education gap, as access to linguistically rich environments can vary widely across different socioeconomic backgrounds.
The data efficiency gap underscores the need for interdisciplinary collaboration between cognitive scientists and AI researchers. By bridging this gap, we can create more sustainable and effective technologies that enhance both our understanding of human cognition and our ability to communicate with intelligent machines. As AI continues to evolve, the lessons from human language learning will be crucial in shaping its future trajectory.
Tags
Original Sources
Kids outlearn AI—and we still don’t know why
↗ https://www.technologyreview.com/2026/08/24/1141740/kids-machines-language-learning
About the author
Amara's entry point into AI was an epidemiology role at a London research hospital, where she spent five years studying how digital health tools reached — or conspicuously failed to reach — underserved communities. Watching early algorithmic systems in healthcare quietly entrench existing inequalities, she redirected her career toward the systemic consequences of AI at scale. She covers AI through an unflinching lens: who benefits, who bears the cost, and what evidence actually says versus what the press release claims. Her writing is calm and precise, but she doesn't mistake balance for neutrality.
More from The Steward →This Week's Edition
31 August 2026
85 articles
Related Articles

Formation Bio Hires Industry Veteran Michael Ehlers as Chief Scientific Officer
Job Market & Society · 3 min

RFK Jr.'s California Tour Highlights Diverse Messaging on Health and Fraud
Job Market & Society · 3 min

Property Tax Hike in Dauphin County Highlights Strain on Public Sector Workers' Health Care Costs
Job Market & Society · 4 min
Related Articles

Formation Bio Hires Industry Veteran Michael Ehlers as Chief Scientific Officer
Job Market & Society · 3 min

RFK Jr.'s California Tour Highlights Diverse Messaging on Health and Fraud
Job Market & Society · 3 min

Property Tax Hike in Dauphin County Highlights Strain on Public Sector Workers' Health Care Costs
Job Market & Society · 4 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.