Why Are Children Better at Learning Languages Than AI? MIT TR: Large Models Need to Burn Down an Entire Forest, While Kids Only Need 100 Million Words

08/26 2026 425

MIT TR In-Depth: Children need only about 100 million words of input to master their native language, while cutting-edge large models require trillions of words for pre-training—how can the data efficiency gap be closed? The answer concerns more efficient AI and understanding the human mind.

It's only been four years since ChatGPT's debut, and we've grown accustomed to 'natural conversations' with our phones and computers. Large language models (LLMs) like Claude, DeepSeek, and GPT are fluent enough to be indistinguishable from humans. But behind the computational curtain lies a stark embarrassment: teaching machines to use human language requires an 'inhuman' amount of data. A cover-story deep dive in MIT Technology Review crunched the numbers—a child mastering their native language hears around 100 million words; a cutting-edge large model pre-training consumes over 10,000 times that. Stanford cognitive scientist Michael Frank puts it vividly: 'We'd have to burn down an entire forest and scrape every scrap of human knowledge to replicate what happens in a living room in just one year.' This gap is what researchers call the 'data efficiency gap.'

When It Comes to Learning Language, Machines Need to 'Burn Down an Entire Forest'

The numbers reveal just how exaggerated the gap is. Meta's open-weight model, Llama 3.1, consumed 15 trillion tokens (word-like units) in pre-training just two years ago; Georgetown University cognitive scientist Wilcox estimates that cutting-edge models may multiply this by ten. But humans? A linguistically rich teenager hears around 100 million words by adolescence; with literacy, that reaches 300 million by age 20. Wilcox draws a comparison: 'The language Claude has seen is equivalent to the sum of experiences of an entire city's generation.' If all the words used to train LLMs were printed on paper, they'd stack beyond the International Space Station; the teenager's 100 million words would only reach 20 meters high. Even more jarring: training GPT-2 with 30 million words yields a 'gibberish generator,' not a child. Frank calls it 'nothing short of a miracle.'

BabyLM Experiment: Can 1 Million Words Approach Large Models?

Image Source: MIT Technology Review

Scientists haven't given up; instead, they're using 'how children learn' as the key to breaking AI bottlenecks. An annual competition called BabyLM requires researchers to train models using only 'developmentally plausible' corpora—around 100 million words (just 10 million for the toddler group), sourced from children's books, conversations, movie subtitles, and real child-directed speech transcriptions—and then test them with grammar questions psychologists use on humans. The results were surprising: the 2024 champion model, GPT-BERT, pre-trained on about 100 million words, outperformed Meta's Llama 2 70B—trained on about 1.5 trillion words—on a BabyLM benchmark. Princeton's Lake even used 61 hours of infant head-mounted camera data to teach models to map words like 'ball' and 'cat' to objects in scenes, all without any innate biases. In other words, language learning can begin with far less data than many theories assume—provided machines 'learn like children' by actively exploring the world through their senses.

The Data Efficiency Gap Hides AI Sovereignty and Computational Power Accounts

Closing this gap matters far beyond academic curiosity. First, there's the computational power and energy consumption account: the internet's 'easily accessible data' will eventually run dry, with some estimates suggesting it could happen as early as the 2030s. Whoever develops 'small-data efficient models' first will burn fewer GPUs and consume less electricity. Second, there's the AI sovereignty and fairness account: Samuel from the University of Oslo points out that languages like Czech, Norwegian, and Sami have only tens of millions of words of trainable data—roughly equivalent to a toddler's exposure. 'How to make small-language models as strong as large-language ones' directly relates to linguistic diversity and national AI autonomy. Third, it challenges the true boundaries of AI education: children still outperform AI in language learning, reminding us that generative AI still has a 'bubble' aspect. For readers aged 36–60 concerned with industry and national fortune, the takeaway is clear: the next phase of U.S.-China AI competition won't just be about data scale and computational power—it'll be about 'who can make models smarter with less.' Data efficiency is the overlooked trump card.

Today's Golden Quote

'The next phase of U.S.-China AI competition won't just be about data scale and computational power—it'll be about 'who can make models smarter with less.' Data efficiency is the overlooked trump card.'

Follow [Degaoxing Zhiqinglang] for three hardcore tech analyses daily, understanding the industry and national fortune behind the technology.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.