Brief IA

OpenAI: $100 Billion for AI Training Data

🤖 Models & LLM·Tom Levy·

OpenAI: $100 Billion for AI Training Data

OpenAI: $100 Billion for AI Training Data
Key Takeaways
1Andrew Ho, former researcher at OpenAI, highlights the increasing specialization of large language models.
2Along with Adam Hunt, he notes that these models excel in programming but are stagnating in other areas.
3Ho anticipates a massive investment of $100 billion in targeted training data.
💡Why it mattersThe evolution of language models may require colossal investments to maintain their versatility and effectiveness.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI: $100 Billion for AI Training Data

A former OpenAI researcher bets that $100 billion will be invested in training data, as he is skeptical that merely scaling up language models will lead to true generalization capabilities.

The new startup he has launched is dedicated to producing high-quality training data, arguing that existing data sources fail to capture many economically valuable skills that AI systems need to operate reliably.

The company is building specialized datasets for bioinformatics and everyday lab work, supported by researchers from the University of Cambridge and Google DeepMind. They note that current AI systems tend to become more specialized rather than versatile, reaching a ceiling in creative problem-solving.

Andrew Ho left OpenAI after just eight months, convinced that large language models generalize poorly. Even in well-funded areas like programming, he claims their performance is inconsistent.

According to Ho, the main cause is a lack of training data. Most economically significant skills are barely represented in existing datasets.

"Most work is highly contextual and cannot be easily encoded in a gradable environment; even if we can observe a 'golden path' taken by a human that we believe is good, it is difficult to understand if alternative or counterfactual paths yield good or bad results," Ho writes.

He argues that scaling alone will not solve this problem and expects AI labs to spend over $100 billion on targeted data collection in the coming years. Ho is also skeptical about the exorbitant valuations of leading labs like OpenAI or Anthropic, which he views as chronically unprofitable since they must continue to invest increasing sums in new models to stay ahead of cheaper competitors like Qwen or Kimi.

His initial products target two areas. The first concerns datasets for complex scientific analyses in bioinformatics, where even current models like GPT-5.6 Sol achieve only about a 30% success rate, a topic he worked on at OpenAI. The second involves datasets for everyday lab work, such as when researchers submit photos of experiments to AI models for evaluation. Fields like chemistry, materials science, health, and broader knowledge work are expected to follow.

The Latest Models May Become Sharper in Some Areas While Dulling in Others

Cambridge researcher Adam Hunt shares Ho's skepticism. Hunt describes how his own view of language models has evolved from initial optimism to growing pessimism. His argument is that the latest models are not becoming more versatile but more specialized. Capabilities in programming and complex mathematics continue to improve, while areas like language quality and simple logic stagnate or deteriorate, corroborating Ho's observation of uneven performance.

Hunt explains that reinforcement learning works well in a domain like coding because clear reward signals and complete training data exist there. In other areas, such data simply isn't available. The early progress of large language models and their apparent ability to generalize through sheer scale and reasoning was a byproduct of training on a broad corpus of text. It was not a sign of true general understanding.

The Question of Generalization is Far from Resolved

This debate always circles back to the same question: can language models develop capabilities that go beyond reproducing and recombining their training data, especially in areas where outcomes cannot be automatically verified? Scientists do not agree on the answer.

Hunt himself estimates his confidence at only 40% and acknowledges that technical advancements could prove him wrong, for example, if specialized AI models can be combined into something resembling more general intelligence.

In a recent position paper titled "LLMs Cannot Leap," Tom Zahavy from Google DeepMind offers a structural explanation for why models are uneven. Language models are good at deduction and induction but fail at creative abduction, meaning the ability to invent a cause for which no linguistic precedent yet exists. As a possible solution, Zahavy suggests controllable world models through action that allow for counterfactual experiments. LLMs could still play a key role in such advanced systems, even if they reach their limits when working alone.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.