Towards AI Labs: DeepMind Outlines a Market, Initial Scores

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Researchers at DeepMind recommend the creation of a market where scientific ideas and experiments would be matched, with rules and payments tied to validation. In automated laboratories, a recent benchmark already measures the share of execution achievable by models and its cost. Google is making strides in biosafety with a dedicated labeling system for biological sequences and structures. Swarm strategies and American public opinion on regulation complete the picture.
DeepMind Proposes a Market for Experiments and Payment Rules
Researchers at DeepMind raise the question of managing a world where the number of scientists would increase significantly and advocate for a market to propose and execute experiments, in order to align ideas with scarce physical resources. They believe that the development of AI scientists will primarily be limited by material resources and empirical validation, rather than by the generation of plausible ideas. They suggest the early creation of an automated scientific economy designed from the outset for this purpose. The market as presented revolves around four components: a proof-of-ideation stage to authenticate provenance before any public evaluation; an ex-ante evaluation phase where different agents make predictions and wager credits or computation tokens; a brokerage system with fractional licensing to connect idea generators with those who execute them; and finally, a payment linked to validation, which triggers royalty payments if the proposal is indeed verified in the real world. It is also noted that if AI manages to tangibly accelerate the work of human scientists, one should expect an increase in demand throughout the scientific supply chain.
Automated Laboratories: 92 Tasks Tested and Quantified Performance
C5R Corp seeks to measure how far AI systems can operate automated laboratories, from molecule manufacturing to X-rays and pill granule pressing. Its benchmark SciUniverse includes 92 tasks spread across 17 families, including sample preparation, instrument control, protocol adaptation, learning through experiments, facility management, and interpretation of real measurements. Examples cover chemistry (assigning structures from NMR), biology (expressing sfGFP in a cell-free system), and materials science (pressing BaTiO3 granules). In terms of results, Claude Fable 5.1 (xhigh) shows a 45.3% success rate at $40.61 per task, ahead of GPT-5 Astra (xhigh) at 32.5% and $52.37, followed by Claude Opus 5 (xhigh) at 30.5% and $46.31. Increasing scientific capabilities could translate into impact, following "wake-up calls" mentioned for 2026 regarding automated R&D.
Biosafety: Google Introduces Dedicated Labeling with SynthID Bio
Google has developed SynthID Bio, a family of labeling methods designed for synthetic biology to enhance biosafety and scientific integrity. The method varies depending on the data: for sequences, it involves selecting certain amino acids, while for predicted 3D structures, it modifies atomic coordinates, allowing for a reliable detection signal. In laboratory experiments on three targets—VEGF-A, the RBD domain of the SARS-CoV-2 spike protein, and PD-L1—the labeled versions showed performance equivalent to the unlabeled versions in terms of success rate, binding affinity, and diversity of natural sequences. This solution is described as complementing other approaches, such as tracking physical equipment and using classifiers from AI providers.
Swarm Agents: Time Savings, Diminishing Returns, and a New Scale Parameter
Toby Ord describes swarms as a new form of inference scale and finds them useful under time constraints. According to him, compared to single agents, swarms consume more tokens, but their parallelization accelerates execution: a swarm of 4 agents requires about twice as many tokens in total for equivalent performance, while halving the expense per agent and potentially the execution time. However, increasing the number of agents comes with diminishing returns, akin to coordination difficulties described by economists through the so-called "stepping on toes" parameter. Multiplying the number of agents by 10 does not yield as much as 10 times more tokens on a single agent; performance follows more like 10λ, between 3 and 5 times in the cited example, with a deficit that accumulates at scale. It is suggested that swarm scaling could even increase the likelihood of an intelligence explosion. More broadly, beyond computation and data for training and inference budgets optimized by thought chains and tool calls, agents add a new scale parameter; better coordination could amplify these returns.
United States: Majority Support for Public Rules Beyond Voluntary Commitments
A survey from the Center for Shared AI Prosperity indicates that the self-regulation agreement announced by the Trump administration with companies like Anthropic and OpenAI is deemed "not sufficient" by 61% of Americans, from a sample of 2,498, including 53% of Trump voters. Furthermore, 54% of voters believe the government should establish and enforce rules for AI, with 61% among Harris voters and 48% among Trump voters. These results are presented as a sign of public demand for stricter regulation than the current position observed in Washington, an asymmetry described as unstable.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.