Richard Sutton Advocates for AI Learning on Real-World Data

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
While Google has written a check for $10 million for data and software from Spirit Airlines, and OpenAI is seeking proprietary corpora, Richard Sutton argues that relying on synthetic data to train models is a dead end. The researcher advocates for learning from real experiences and launched Oak Lab last month to develop agents that continuously learn from their environment. His comments were made during a podcast on Tuesday.
The Quest for Real Data Intensifies Among AI Giants
Major tech companies are ramping up efforts to access real data that is hard to find online. OpenAI has publicly sought large-scale proprietary datasets to train its models. Earlier this week, Google agreed to pay $10 million to obtain internal data and software belonging to Spirit Airlines, which is currently in bankruptcy.
Experience-Based Learning: Sutton's Thesis Put into Practice by Oak Lab
Richard Sutton emphasizes the importance of real experience data, meaning information collected by an agent interacting directly with its environment, with continuous observation and learning from the consequences. He founded Oak Lab last month in collaboration with Khurram Javed, who was previously his student. The company designs agents whose learning is based on continuous acquisition from their own experiences, rather than relying primarily on large pre-selected datasets. To date, Oak Lab has not provided any information regarding a potential funding round or the identity of its investors.
Why He Rejects Synthetic Data for Training Models
On Tuesday, Richard Sutton described the use of synthetic data to advance large language models as a "big mistake." He argues that manufactured data cannot replace real data, citing human behavior as a counterexample: "There’s no way to have synthetic data about the minds of others." Regarding physical systems, he claims that artificial data cannot predict how a drone will interact with the world due to variables like friction and engine wear, summarizing: "The world is infinitely complex, and any simulation of it is microscopic." Synthetic data, artificially produced by algorithms rather than collected from the real world, is becoming increasingly attractive as companies scour the web, with examples ranging from images of cars for autonomous driving to fake bank statements for fraud detection. In an episode of the Sequoia podcast released on Tuesday, he also expressed that the industry is heading in the wrong direction by relying on these corpora.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.