Anthropic and Stanford: Why Only Large Models Succeed
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A Study That Challenges Conventional Wisdom
A recent study conducted by researchers from Anthropic, Stanford, and other institutions has shed light on a new approach to understanding why large language models acquire skills that smaller models struggle to master. Contrary to popular belief, this is not simply a matter of learning speed. The researchers suggest that instead of indefinitely increasing model size, it may be more effective to increase the frequency of certain specific tasks in the training data to anchor rare skills in smaller models.
The Limitations of Small Models
The study highlights that small models often fail to reliably learn rare tasks, even when training periods are significantly extended. Well-known scaling laws show that a small model never reaches the same loss as a large model, regardless of the amount of data provided. This phenomenon can be explained by the fact that frequent and simple tasks take priority in the learning process, relegating rare and complex tasks to the background.
The Dynamics of Frequent and Rare Tasks
To isolate the mechanism at play, the researchers tested a mix of tasks with varying frequencies and complexities. A model with a given number of neurons is assigned the most useful features, where usefulness is determined by the frequency of a task's occurrence and its importance. In the experiments conducted, only sufficiently large models learned tasks that constituted only 0.25% of the training data. This is because as long as frequent tasks are not well mastered, they strongly pull the model in their direction at each training step, thereby overshadowing much of what the model has learned about rare tasks.
Experiments and Observations
The researchers designed an experiment to clearly separate this effect. The total frequency of a rare task remains constant, but the gap between individual observations varies. The larger this gap, the more the signal degrades in narrow models. In contrast, wide models better retain the signal between observations and rely on it. To test this theory during pre-training, the team trained OLMo models ranging from 4 million to 4 billion parameters on up to 210 billion tokens from the Dolma corpus. They mixed two artificial tasks in the data, a number comparison and a modular addition, with frequencies ranging from about 1,000 instances per batch to one instance every ten batches.
The Results of Real Language Models
In the results obtained, only the large OLMo models learned the rare tasks by understanding the underlying rule and applying it to new cases, rather than simply memorizing individual examples. This was particularly clear with modular addition, where the researchers observed what is known as "grokking." A model first memorizes a task, then suddenly understands the actual principle after additional training. Only the largest models reach this moment, and only when the task appears frequently enough in the data.
Memorization as an Intermediate Step
The study considers memorization as a prerequisite for generalization, rather than an undesirable side effect. A model must retain individual observations long enough for a broader pattern to form across many batches. This offers a practical alternative to simply increasing model size. Instead of increasing the model size, the frequency of a target task in the training data can be increased to anchor a specific skill, the research suggests.
Towards a New Training Strategy
There are several theories explaining why model size helps. In May, a team from MIT linked scaling laws to the geometry of models, where models store more concepts through overlap than their dimensions should allow. This new study takes a different angle, focusing on what a model can actually learn from a given mix of data during training. The older debate about whether capabilities "emerge" during sudden jumps beyond a certain size, or if this is partly a measurement artifact, is still ongoing.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.