Large language models (LLMs) require vastly more data than children to learn language. Understanding why could help us create more efficient models and reveal more about developing minds.
For at least 100,000 years, human language acquisition has been a uniquely biological marvel, with human children mastering complex linguistic systems with astonishing efficiency. Now, a new contender has emerged: artificial intelligence. In just four years, LLMs like Claude, DeepSeek, and OpenAI’s GPT models have achieved remarkable fluency, capable of engaging in natural, flexible conversations that can convincingly mimic human interaction. However, this impressive feat comes at a significant cost. Teaching these AI models to use human language demands an inhuman volume of data, consuming hundreds of thousands of times more words than a child experiences while mastering their native tongue, and far more than a child typically hears by their first birthday, when their language acquisition journey truly begins.
"The progress recently has been amazing," acknowledges Michael C. Frank, a cognitive scientist at Stanford University, referring to LLMs. "But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year." This stark contrast, known as the data efficiency gap, highlights a fundamental question for cognitive scientists and AI researchers: How do children consistently outperform even the most sophisticated AI models in linguistic learning?
Bridging this gap holds significant implications for both AI development and cognitive science. For a decade, AI language models have primarily advanced by increasing their size and the data they consume. Meta’s Llama 3.1, for instance, underwent pretraining on 15 trillion tokens (word-like units of language), and current frontier models may be trained on ten times that amount. This reliance on ever-increasing datasets raises concerns about future scalability, as the readily available pool of internet data could be depleted as early as the 2030s. Children, however, demonstrate that learning can be achieved with significantly less data. A preteen in a linguistically rich environment might hear around 100 million words, a figure that could rise to 300 million by age 20 with the addition of literacy. The sheer scale of data used to train LLMs is staggering; printing it all would create a stack reaching beyond the International Space Station, while a child’s 100 million words would stack a mere 20 meters high.
By reverse-engineering the mechanisms of child language acquisition, scientists aim to develop more data-efficient AI models. This could unlock advancements in areas like training AI on video data and creating chatbots for minority language communities. Furthermore, testing hypotheses about human learning within AI models can help resolve long-standing questions about language and the development of the human mind. Are we born with an innate language faculty, or is language acquisition purely environmental? Is our language processing a biological quirk, or does it reflect universal principles of language use and learning?
The complexity of language becomes apparent only when attempting to learn a new language as an adult, grappling with tenses, sounds, grammatical cases, and gendered nouns. Yet, mastering one’s mother tongue is typically effortless. Toddlers often begin producing grammatically correct sentences after hearing approximately 10 million words, with some estimates reaching up to 30 million. "It’s just totally miraculous," states Frank. "If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid."
The precise methods by which infants achieve this are still a mystery. While researchers have gained considerable knowledge about the stages of child language development, fundamental questions remain unanswered. Chief among these is why babies can learn language at all. Human language syntax, the rules for constructing sentences, involves recursive, nested structures that enable the expression of virtually infinite ideas using a finite vocabulary. This complexity should pose a significant challenge for infants, who are exposed to only a fraction of the language they will eventually master. Yet, from this limited exposure, they infer the entirety of the linguistic system.
One prominent explanation, proposed by linguist Noam Chomsky in the 1950s, suggests that infants are born with an innate, hardwired knowledge of grammar. This countered the prevailing behaviorist view of psychologist B.F. Skinner, who argued that language is learned solely through environmental conditioning and reinforcement, akin to training a dog. Chomsky’s "poverty of the stimulus" argument posited that the complexity of language, particularly syntax, is too intricate and children’s exposure to it too "impoverished" to be learned entirely from experience. "His signature argument was, essentially, that language cannot be learned on the basis purely of statistics," explains Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Chomsky theorized that language is based on a set of logical rules, and children must possess innate knowledge of these rules to deduce grammar from fragmented speech.
The Chomskyan perspective, known as generative grammar, dominated linguistics in the United States for decades and significantly influenced early AI research. During the initial boom of AI in the 1950s and 60s, fueled by military funding aimed at developing language translation capabilities, the lines between linguistics and natural-language processing blurred. Despite early successes with simple neural networks that could recognize and reproduce statistical patterns, many US AI researchers adopted a rule-based framework influenced by Chomsky. This approach, known as symbolic AI, involved explicitly coding linguistic rules into programs, essentially treating language learning like a grammar class rather than an immersive experience. This approach prevailed for decades but largely failed to produce models capable of handling human language at scale, contributing to the "AI winter" that began in the 1970s.

Neural networks began a resurgence in the aftermath of the AI winter. By the 2010s, advancements in affordable and powerful computer hardware, coupled with the exponential growth of the internet, led to a dramatic improvement in their performance. By 2018 and 2019, models like BERT and GPT-2, built on the transformer architecture and trained on billions of tokens, demonstrated the efficacy of learning from massive datasets for language. The breakout success of OpenAI’s ChatGPT in 2022 made this capability evident to the wider public.
LLMs are not biological brains; they are sophisticated statistical learners. These "naïve pattern-learning machines," devoid of the evolved biological nuances of the human cortex, are precisely the type of system that generative linguists once believed incapable of learning language. Yet, they now produce convincing prose and pass grammar tests. "No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax," remarks Alison Gopnik, a developmental psychologist at the University of California, Berkeley. "I didn’t think that was going to turn out to be true. And I think most people didn’t think that you could just look at the statistics of a large sample of language and figure out grammar."
The critical question remains: Can models learn effectively from a small sample of language, comparable to a child’s exposure? Can we create "baby-scale" models that go beyond mere nonsense generation?
Alex Warstadt, a linguist and data scientist at the University of California, San Diego, recalls the transformative period surrounding the release of BERT and GPT-2. In 2019, as a PhD student, he witnessed his field undergo a profound shift. The mere ability of language models to learn English from text challenged prevailing Chomskyan theories, though many linguists remained unconvinced that LLMs could offer insights into human language acquisition. "I always got pushback on one issue in particular. And that was the size of the data sets of the model," Warstadt notes. "There was never a time when people were training language models at human scale where we were impressed by them."
Warstadt, however, saw potential in LLMs. He reasoned that even imperfect scientific models can be informative, and LLMs represented powerful simulations of human language use. By embedding hypotheses about child learning into these models and measuring their performance in closing the data gap, scientists could potentially test their theories. In August 2022, Warstadt initiated a discussion on Twitter, arguing for the utility of neural networks as models of language acquisition. Following interactions with AI researcher Leshem Choshen, the idea for BabyLM, an annual competition focused on training models on limited datasets, began to take shape.
Four years later, BabyLM has expanded to include workshops and spin-off competitions, such as one for baby models trained on Chinese. The core challenge involves training language models on a "developmentally plausible" corpus of just 100 million words (or 10 million for a toddler-scale track), drawn from diverse sources like storybooks, dialogues, movie subtitles, and various versions of Wikipedia. The models are then evaluated using psycholinguistic benchmarks designed for human subjects, according to Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University and one of the competition’s organizers.
One evaluation method involves presenting both human and machine participants with sentences and measuring their reactions to ungrammatical elements. For example, a test might compare "The keys to the cabinet are on the table" with "The keys to the cabinet is on the table." Wilcox explains, "When humans see ‘is,’ they’re like: What? That’s not supposed to be ‘is’…" This "surprise" in humans can be measured through eye-tracking, while in language models, it’s quantified by "surprisal," a metric assessing the model’s predicted likelihood of a sentence or phrase.
The competition has already challenged some long-held assumptions, such as the effectiveness of "curriculum learning"—a method that starts with simple training data and gradually introduces more complex inputs, mirroring how humans might learn through stages. This approach was the most popular in the first BabyLM round but did not yield the expected results. "The appeal is just kind of hard to resist, you know. [Curriculum learning] seems to really line up with ways that we believe humans are learning," says Aaron Mueller, a computer scientist at Boston University and a BabyLM organizer. "But it seems like these transformers don’t really need to have their data ordered in such a way to learn effectively."
Ironically, the top-performing BabyLM models are not directly inspired by infants. The 2024 champion, GPT-BERT, is a transformer that combines predicting the next token in a sequence (like modern LLMs) with BERT’s "masked language model" approach, which involves filling in missing tokens. Remarkably, when pre-trained on approximately 100 million words, GPT-BERT outperformed Meta’s Llama 2 70B—an LLM pre-trained on roughly 15,000 times that amount—on one of the BabyLM benchmarks.

Despite these advancements, BabyLM models still lag behind commercial LLMs. Many cannot generate text, and even GPT-BERT would appear rudimentary next to a contemporary AI chatbot. Crucially, the learning mechanisms of these "baby" models are not inherently child-like. Children are not disembodied programs whose sole "experience" of the world comes from written text. They actively engage with their environment through their senses, particularly vision and hearing. To truly close the data gap, some researchers believe, machines must learn through the sensory experiences of children.
When Michael Frank established his lab at Stanford about 15 years ago, the understanding of how babies experience the world was still developing. Developmental psychologists were beginning to gain insights through head-mounted cameras. "The insights that came out from that early research were that kids’ experience looks really radically different than we thought," Frank observes. "It’s much more focused: They’ve got these little short arms, so the objects are, like, right in front of them. And they live in a forest of knees."
Frank, eager to use headcam footage to test hypotheses about child language learning, required more data. He and four colleagues recruited three infants—children of psychologist mothers who understood the project’s aims—to wear headcams. The SAYCam project captured two hours of each child’s life per week, from six months to two and a half years of age. "The families were willing to release that video, and that’s critical," Frank emphasizes. "So we released it, and people started training models on it."
Brenden Lake, a cognitive scientist and AI researcher at Princeton, was one of those individuals. In 2024, while at New York University, he and his colleagues presented a model trained on 61 hours of SAYCam data that successfully identified objects and associated them with words. Many developmental psychology theories propose that infants possess innate biases to help them segment sensory input and link it to language. For instance, it’s believed babies assume a new word like "shoe" refers to the whole object rather than a part (like a shoelace). However, Lake’s model learned to identify objects and associate them with words without such biases. "It turns out you can get a real start on language learning using a lot less than what a number of theories suggested," Lake states. Nevertheless, he adds, "we don’t get a two-year-old out of [training] when we’re done."
The limitations of current models might stem from the insufficient realism of the training data. SAYCam and its successor, BabyView, capture only a few hours of footage per week. Researchers face a choice between analyzing a small segment of a single child’s life or pooling data from multiple children, neither of which fully replicates a child’s lived experience.
This landscape may be changing. Uri Hasson, a neuroscientist and psychologist at Princeton, has spent the past five years leading a project to record the first 1,000 days of 17 children’s lives. Participating families equipped their homes (excluding bedrooms and bathrooms) with cameras and microphones, recording 12 hours daily. This unprecedented dataset, detailed in a recent preprint, would be unmanageable without advanced AI tools for transcription and video analysis, according to Hasson. "For the first time, we have the input," he states. "It’s really only the beginning."
Training models on video data has proven challenging. While text-based models achieve fluency after ingesting vast datasets, multimodal models trained on child-centric video data are far from comparable. Lake’s model, for example, learned simple words, but attempts to integrate visual data into BabyLM have not yielded similar results. Gopnik suggests that children are not passive observers; they are active explorers who "actively choosing their own data." This constant experimentation might be the missing ingredient.
Research from Gopnik’s group, including studies of schoolchildren exploring a Minecraft-inspired game, demonstrates that play is an effective method for learning cause and effect. Children seek experiences and take actions that maximize their "empowerment"—their ability to predictably influence their environment.
Unlike AI models, children possess an awareness of their knowledge gaps and a drive to fill them, according to Elizabeth Bonawitz, a developmental cognitive scientist at Harvard. Furthermore, children’s social interactions play a crucial role in their learning. Bonawitz’s research indicates that children interpret information differently when they perceive an adult is actively teaching them. "Children are not only reasoning about the evidence they’re being told," Bonawitz explains. "They’re reasoning about the teacher, about the teacher’s knowledge, and about why the teacher is telling [them] this particular information."

This contrasts sharply with the passive, isolated learning of AI models. If models were designed to actively seek information to address their blind spots, experiment with language, observe reactions from other language users, and reason within a simulated social context, their learning might improve. Last year’s BabyLM competition even included a category for models that could learn through interaction with other models, but these social models did not outperform standard ones.
Among leading industry labs, Meta appears most inclined to draw inspiration from children, particularly for video-based model training. Two Meta researchers contributed to BabyLM’s multimodal branch, and Meta scientists, collaborating with academic researchers like Frank, recently introduced a benchmark and challenge for training models on baby headcam footage. Frank also noted interest from Flapping Airplanes, a stealth-mode AI startup. None of these companies—Meta, Google DeepMind, OpenAI, or Flapping Airplanes—agreed to an interview.
For now, Gopnik believes that frontier labs are not prioritizing the adoption of child-like learning strategies. She anticipates that the next generation of AI, beyond the current transformer architecture, will be the first to significantly integrate lessons from developmental psychology.
"Perhaps the most enticing reason to close the data gap is that it could help us understand ourselves."
Mueller observes that the machine learning community generally prioritizes functional performance over mimicking the brain. However, he notes a growing awareness and interest in the data efficiency gap, citing the NanoGPT Slowrun benchmark launched by Q Labs in March 2026, which shares similar goals with BabyLM but focuses solely on data efficiency rather than human language learning.
Warstadt’s motivation for closing the data gap extends to democratizing AI, enabling universities and smaller organizations to train competitive models without massive resources. David Samuel, a machine learning researcher at the University of Oslo and a co-creator of GPT-BERT, has a more personal reason: as a Czech speaker working in Norway, he recognizes the scarcity of Czech and Norwegian data for LLM training compared to English. Minority languages like Sami may have only tens of millions of tokens available—a scale comparable to a toddler’s exposure. "The question was," Samuel states, "how can we develop language models that are just as capable as the English ones for small languages?"
However, the most compelling reason to bridge the data gap may be its potential to deepen our understanding of human cognition. Bonawitz initially doubted the relevance of LLMs to cognition, given the fundamental differences between AI and brains: brains are embodied, neurons are living cells, and they mature and adapt throughout life, unlike pre-trained LLMs. Yet, she admits, "I’m sort of revising my beliefs." She now sees value in studying models as a form of comparative psychology, akin to studying animal minds to illuminate our own.
Researchers like Warstadt, Frank, Wilcox, Lake, and Hasson are already employing language models as "linguistic lab rats"—imperfect yet informative stand-ins for human language users. This approach is particularly valuable for questions concerning learning, language, and information processing that are not strictly tied to brain biology. When models achieve feats previously thought impossible, it challenges long-held assumptions. Researchers can systematically embed hypotheses about language learning into models—simulating varying degrees of bilingualism or withholding exposure to specific grammatical forms—and test these hypotheses in ways impossible with human children. Futrell likens this to teaching language to an alien and then examining its internal workings.
While other animals communicate, only humans converse. Now, a non-animal, non-human entity can also engage in discourse. LLMs offer opportunities for comparative studies, even with the profound differences between models and minds. "For the last 100,000 years or however long human language has existed, humans have been the only entities in the universe that use language. Now there’s this other linguistic entity," Warstadt remarks. "Finally we have a model; not in the sense of a language model, but in the sense of a model organism."

