Learning Like a Child: AI Struggles with Data Efficiency

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Language models have learned to handle syntax and persuade in writing, but at the cost of an unparalleled appetite for data. Children, on the other hand, acquire grammar with minuscule exposure in comparison. AI and cognitive science researchers are exploring how to replicate this efficiency, as online corpora may reach a limit as early as the 2030s.
Systems that grasp syntax, but the "baby" scale remains a hurdle
Large language models write plausible sonnets and pass grammar tests. Alison Gopnik notes that these systems learn syntax, a result that she and others did not think achievable by exploiting statistics on vast corpora. A concrete question remains: can a model trained with exposure at a baby’s level produce anything other than nonsense? Alex Warstadt refers to the period around BERT and GPT-2 as an exciting turning point, as in 2019, he observed, as a doctoral student in New York, his field shifting. However, many linguists remained skeptical about the contribution of LLMs to the study of acquisition, mainly criticizing the size of the datasets: at a human scale, no model impressed. Nevertheless, Warstadt defends their utility as powerful simulations of language use and proposes integrating hypotheses about children's learning, then measuring performance to test these ideas. He publicly articulated this argument in August 2022. Meanwhile, for the general public, naturally conversing with assistants has become commonplace, and models like Claude, DeepSeek, or GPT hold their own in written comparisons with humans.
From a symbolic bet to the transformer era, through the AI winter
In the mid-20th century, generative grammar fueled computer science, as military funding pushed to equip machines with capabilities for understanding English and translating Russian. Despite early successes of simple neural networks, the rule-based approach dominated for decades, failing to produce systems capable of handling language at scale, leading to the "AI winter" of the 1970s. Starting in the 2010s, with more affordable hardware and a broader Internet, neural networks regained control. In 2018 and 2019, BERT and GPT-2, built on the transformer architecture and trained on billions of tokens, demonstrated that a massive surplus of data allows for concrete advancements. In 2022, this reality burst into the open with ChatGPT. These models remain powerful statistical learners, not brains, something many generative linguists would not have deemed possible two decades earlier.
The major constraint: data that is not infinite
The last decade has primarily reduced error by enlarging models. Llama 3.1, with open weights, was pre-trained on 15 trillion tokens, and some estimate that cutting-edge systems could still multiply these volumes by ten during pre-training, the preliminary step before fine-tuning for a task. However, the Internet does not contain infinite data, and the easily exploitable source could dry up as early as the 2030s. Teaching language to machines today requires truly inhuman corpora: an LLM easily ingests one hundred thousand times more words than a speaker processes for their native language, and far more than what a baby hears before their first birthday. Michael C. Frank praises "incredible" progress while highlighting the disproportionate informational cost required to replicate in the lab what happens in children within a year. This is the crux of the data efficiency gap that challenges both cognitive researchers and model architects.
What children prove: minimal exposure is often sufficient
A preadolescent, even in a very rich context, may have heard around 100 million words; with reading and writing, total exposure can peak at around 300 million by age 20. Yet, toddlers typically begin producing grammatically correct sentences after hearing about 10 million words, or even 30 million in favorable cases. Michael C. Frank observes that a GPT-2 limited to 30 million words produces only nonsense, nothing approaching a child's output. In contrast, modern models would have seen the equivalent of an entire generation of language from a city, according to Ethan Gotlieb Wilcox. On paper, their training corpus would reach the International Space Station, while the 100 million words of a preadolescent would form only a 20-meter stack—and humans can manage with even less. The species has been speaking for at least 100,000 years, but until very recently, only a human child accessed a human language with such ease.
Tension in theories and promises of a more data-efficient AI
While many milestones of acquisition can be described, how babies achieve this remains a mystery. Human syntax combines recursion and embedding to express virtually an infinite number of ideas with a finite lexicon, even as the sample heard by babies remains limited. In the 1950s, Noam Chomsky defended the idea of innate grammatical knowledge in response to B.F. Skinner, who viewed language as a conditioned learning process. The poverty of the stimulus argument asserts that statistics alone are insufficient, as Richard Futrell summarizes, and posits underlying logical rules necessary for the child to deduce grammar. This view dominated American linguistics under generative grammar. The fact that models learn English from text has challenged these ideas, yet it does not explain infant efficiency.
To move forward, teams are planning to reverse-engineer children's learning to build more data-frugal models. Such advancements could aid training on video or the design of chatbots for minority language communities. Testing these hypotheses in models could also provide answers to enduring questions: is there a linguistic instinct, can learning rely solely on experience, and what portion is attributed to biology or more universal constraints on languages?
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.