Children outpace AI in language learning
Reported by MIT Technology Review AI: Kids outlearn AI—and we still don’t know why. Analysis and context written by TickrWire.
A new analysis highlights the stark data gap between children and large language models, showing that children master language with far fewer exposures. Researchers are using the BabyLM competition to probe how limited data can still yield linguistic competence.

- Children can master language after hearing roughly 100 million words, while top LLMs need 15 trillion tokens.
- The BabyLM competition trains models on just 100 million words and has already produced models that beat much larger systems on grammar tasks.
- Curriculum learning, which orders data from simple to complex, did not significantly improve performance in the BabyLM tests.
- Head‑camera studies show that multimodal visual‑audio exposure can help models learn word‑object links with far less text.
- Experts warn that the pool of publicly available training text may run out within decades, prompting a search for more efficient learning methods.
When Michael C. Frank, a cognitive scientist at Stanford, declared that large language models have made ‘amazing’ progress, he was pointing to a yawning divide between artificial systems and human children. For decades, the only known entity that could learn a human language to perfect fluency was a child raised in a home. The arrival of ChatGPT four years ago gave the impression that machines could converse as naturally as people, yet beneath the fluency lies a stark data inefficiency. Children acquire language from a modest stream of speech, while today’s biggest models consume astronomical amounts of text. This contrast forms the core of a recent Technology Review feature that investigates why children still outperform the most sophisticated AI ever built.
The magnitude of the difference is difficult to grasp without numbers. A preteen raised in a linguistically rich household may have heard roughly 100 million words, and adding literacy by age 20 can push that total to about 300 million words. In contrast, Meta’s Llama 3.1, released two years ago, was pretrained on 15 trillion tokens, word‑like chunks that dwarf the child’s exposure many times over. Ethan Gotlieb Wilcox of Georgetown University notes that frontier models could be trained on ten times more data, but the internet’s supply of easily accessible text is finite. The article likens the LLM’s data diet to a stack of paper that would reach past the International Space Station, while the child’s 100 million words would form a pile just 20 metres high.
For the past decade, language models have mostly improved by scaling up parameters and ingesting ever larger corpora. Llama 3.1’s 15 trillion‑token pretraining run exemplifies this strategy, and industry analysts expect future models to require even more data. Yet experts warn that the well of publicly available text may run dry as early as the 2030s, creating a bottleneck for continued performance gains. The scarcity of training material forces researchers to reconsider whether sheer volume is the only path forward, or whether insights from human learning could unlock more efficient alternatives.
In response, a growing number of labs have turned to the BabyLM competition, an annual challenge organised by Warstadt, Choshen and other researchers. The competition asks teams to train language models on a ‘developmentally plausible’ corpus of just 100 million words (or 10 million for a toddler track) drawn from storybooks, dialogue, movie subtitles, Simple English Wikipedia, normal Wikipedia and actual child‑directed speech. The 2024 champion, GPT‑BERT, combined a next‑token prediction objective with a masked‑model loss, and when pretrained on roughly 100 million words it outperformed Meta’s Llama 2 70B, a model trained on about 15 trillion tokens, on a grammar benchmark. The result suggests that architectural tricks and objective design can partially compensate for limited data, though GPT‑BERT still falls far short of commercial LLMs in overall fluency.
Organisers also tested curriculum learning, which presents simple data before moving to more complex inputs, mirroring the presumed progression from baby talk to adult language. Aaron Mueller of Boston University, a BabyLM organiser, admits the approach ‘seems to really line up with ways that we believe humans are learning’, but the transformers in the competition did not benefit significantly from such ordering. Warstadt observed that the models learned effectively regardless of data schedule, indicating that the ordering hypothesis may not be essential for these artificial networks. The failure of curriculum learning to boost performance highlights how little is still known about aligning artificial training regimes with the subtle cues that guide child development.
Beyond text‑only experiments, researchers such as Brenden Lake have begun using head‑camera footage from the SAYCam project, which recorded two hours a week of three infants’ lives from six months to two-and‑a‑half years. The model Lake’s team built on 61 hours of raw video learned to associate visual objects with spoken words without pre‑programmed biases about word‑object mapping. This work suggests that multimodal exposure, seeing the referent while hearing the label, could compress the data needed for language acquisition.
Looking ahead, the central question remains whether AI can ever match the data efficiency of a human child. The poverty‑of‑the‑stimulus debate, originally championed by Noam Chomsky, still frames discussions about innate linguistic knowledge versus pure statistical learning. If models can be built to leverage limited, structured data, whether through architectural priors, multimodal inputs, or curriculum‑aware training, then the field may see chatbots that serve minority language communities or video‑based assistants that require far less footage. What to watch in the coming years includes the expansion of BabyLM to new languages, the emergence of head‑cam‑derived datasets, and any breakthroughs that claim to close the data gap without resorting to ever‑larger corpora.
May inspire architectures that learn from limited labeled data.
Could lower compute expenses for language‑based products.
Signals a potential ceiling in data‑driven scaling of language models.
Provides a real‑world test case for theories of language acquisition.
Illustrates how humans outperform AI on a core task with far less input.
- tokens
- word‑like chunks used by language models to process text.
- surprisal
- a measure of how unlikely a model predicts a given word or sentence.
- data efficiency gap
- the difference in data quantity needed for children versus AI to achieve language fluency.
AI ResearchPew study confirms sharp rise of AI-written text on the web since ChatGPT's launch
AI ResearchWho’s behind the new ‘stealth model’ Ox Alpha?
AI ResearchAI could make scientists do more work less well, not less work better, study argues
AI ResearchStudy explains why AI agents benefit from "skills" and when they fail
AI ResearchWorld models that ignore human beliefs predict the wrong actions, new research shows
BusinessTrump bought SpaceX shares two weeks after blockbuster IPO
President Donald Trump purchased up to fifty thousand dollars in SpaceX stock shortly after the company's initial public offering, according to financial disclosures.
BusinessAmjad Masad, CEO and co-founder of Replit, joins the Disrupt Stage at TechCrunch Disrupt 2026
Replit co-founder and CEO Amjad Masad is scheduled to speak at TechCrunch Disrupt 2026, discussing the evolving software landscape and his company's rapid financial ascent amid the artificial intelligence boom.
SecurityInstinct’s powerful AI assistant is raising privacy and security concerns
Instinct, a new AI personal assistant, is drawing attention for its powerful features but also raising serious concerns about user privacy, security, and control over personal data.
HardwareCerebras unveils CS-4 with double the performance on the same chip
Cerebras has launched the CS-4, a rack-scale AI accelerator that doubles performance over its predecessor by optimizing power and cooling for the WSE-3 chip.
SecurityFlock CEO calls for ‘compromise’ as surveillance company faces growing backlash
Flock Safety CEO Garrett Langley is advocating for a compromise between public safety and privacy as the surveillance tech company encounters intense scrutiny and political pushback over alleged misuse.
SecurityIs it legal to train AI models on copyrighted books? It’s complicated
Courts are split on whether training AI on copyrighted books counts as illegal copying or protected fair use, leaving authors and tech firms in legal limbo.