Faraday Surpasses 73% in Science, Two AIs Reach Level 7 of DiG-bench

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
An AI model for science, Faraday, announces superior performance of over 73% on replication tasks and surpasses generalist models in certain cases. Meanwhile, DiG-bench, a benchmark of 70 text games with hidden rules, clearly distinguishes systems capable of discovery, with only two models completing level 7 tasks. A simulator for research on recursive self-improvement and a manifesto from Mark Zuckerberg complete the current landscape of ambitions and tools.
Faraday surpasses generalist LLMs on replications
Researchers from the startup Inherent describe the construction of Faraday, a scientific AI model. The team established a supervision harness and a relatively small LLM driving large proprietary models to enhance their efficiency in science. To evaluate it, they assembled a set of 100 AI/ML articles published between 1990 and 2026, converted into 310 replication tasks. Faraday, using Codex, outperformed standard models like Opus 4.8 and GPT-5.5 on some of these tasks. On science-related AI tasks, it achieves performance exceeding 73%.
DiG-bench evaluates discovery in 70 hidden-rule games
DiG-bench offers 70 text games, designed by human experts, where rules and objectives are hidden and must be discovered through interaction. Researchers from institutions including Oxford, Princeton, KAUST, the Swiss AI Lab, Inria, and MIT have created a framework aimed at measuring the ability to identify the mechanisms that determine a player's success. The majority of the games remain private, but 21 are public; most feature multiple levels and offer between 2 to 34 possible actions, with an optional mode allowing unlimited actions. All games have been beaten by at least one human and are often deemed challenging, requiring varied skills and strategies.
The benchmark is structured into seven increasing levels of difficulty. Opus 5 and Fable 5 with Claude Code lead the pack, being the only ones to succeed in level 7 tasks; Opus 5, GPT-5.5, and Kimi K3 validate certain level 6 tasks. The games are presented as sufficiently demanding to still elude current models, although some achieve notable discoveries. The overall aim is to isolate a prerequisite for creativity, namely the autonomous discovery of useful elements in new situations. In this context, Fable is also described as demonstrating a form of creative intuition.
A simulator explores the trade-off between computation, researchers, and data
Paradigm Research provides a game that simulates the management of a company working on AI systems with recursive self-improvement capabilities. This tool allows for an understanding of how various aspects of AI research are interrelated, particularly the allocation of investments between researchers and computing power, as well as decisions regarding the use and timing of data exploitation.
Zuckerberg outlines his goals for agents and tutoring
Mark Zuckerberg publishes a manifesto-essay, "The Future is for Everyone," which outlines his approach to developing AI systems at Meta. He advocates for a philosophy of individual empowerment as a source of prosperity and positions invention as the ultimate goal of superintelligence. His stated objectives include a highly capable personal agent for everyone, tools for idea generation, tools for creating new businesses, and a personalized tutor with a PhD level in every subject.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.