Brief IA

DeepMind Expands Co-Scientist: Measured Results and Limitations

🤖 Models & LLM·Tom Levy·

DeepMind Expands Co-Scientist: Measured Results and Limitations

DeepMind Expands Co-Scientist: Measured Results and Limitations
Key Takeaways
1With its active modules, Co-Scientifique produced key results in only 4% of cases, compared to 46% without them and 90% for a comparison system.
2In materials, three thin films were synthesized on the first attempt, and in biology, three out of four characteristics were correctly predicted.
3The advantage of Agent_H on benchmarks was only reflected in one category during an evaluation by three doctors.
💡Why it mattersThe system now covers the entire research cycle with measured safeguards, but limits confirmed by human evaluations and transfer uncertainties constrain its use.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google DeepMind's multi-agent system has transitioned from simple hypothesis reasoning to a complete research cycle, from experimental planning to article writing. Validations exist in three disciplines, but gaps between benchmarks and clinical opinions, transfer uncertainties, and residual errors strictly frame the scope of the results.

Gaps with Human Opinion and Methodological Reservations

The performance displayed by Agent_H on health benchmarks did not hold up during an evaluation by physicians. Three certified practitioners rated the responses across nine categories and observed a statistically significant advantage over Gemini 3.1 Pro in only one dimension: the reduction of the risk of potentially harmful responses. Furthermore, automatic evaluations correlated poorly with clinical judgments, and researchers remind us that high scores do not imply better responses in practice. Beyond the benchmarks, residual errors persist: the system tends toward selective reporting and may describe plausible methods that do not correspond to its code, according to Samuel Schmidgall. In materials science, after 25 iterations, structures close to the target were obtained, but atomic confirmation remains pending, and the transferability of recipes to other laboratories remains uncertain.

What the System Now Achieves in the Laboratory

Co-Scientist relies on Gemini to plan experiments, program and control equipment, then analyze results and generate manuscripts. Its project management follows three explicit steps: ideation, experimentation, and writing. The application scope covers scenarios where humans perform syntheses, collaborations in biology, and AI architecture developments conducted autonomously.

Three Demonstrations: Materials, Biology, Computer Science

According to Google, validated results have been achieved in three disciplines. In the field of materials, teams connected the system to a semi-automated furnace, allowing it to suggest a safer method for obtaining a 2D material that had previously been primarily manufactured using methods deemed hazardous, by adjusting protocols to the laboratory's equipment. In another experiment, the synthesis of three thin layers of semiconductors was successfully achieved on the first attempt. The direct use of Gemini 3 Deep Think reduced the time required to develop protocols from several days to just a few minutes, although the placement of samples and precursors continued to be done manually, and the accelerated option produced smaller and less uniform crystals. In biology, the system independently built an image analysis pipeline to predict E. coli colony patterns based on chemical concentration; predictions from Gemini 3 Pro Image coincided with unpublished laboratory results for three out of four characteristics, while remaining limited to interpolation between known conditions. In computer science, a fully autonomous run after initial configuration led to the design of Agent_H, which classifies queries, generates dozens of candidate responses in parallel, and then refines them; after correcting the output length, Agent_H surpassed six benchmark models on health benchmarks, including GPT-5 and Claude Opus 5.

Reliability: 4% Fabrications with Active Safeguards

The tendency of autonomous systems to fabricate information is documented, with analyses reporting rates of 80 to 100% in some cases. Co-Scientist combines penalties against fabricated or plagiarized content and a separate verification that recalibrates each numerical value against the results of the executed code. Tested in a double-blind study on 150 articles by 30 experts for 450 evaluations, this architecture reduced key fabrications to 4% when the modules were active, compared to 46% without them and 90% for a comparison system. No entirely invented data was observed in Co-Scientist's results, while it appeared in 44% of the comparison system's articles. Near-plagiarism decreased from 60% to 16%, and the security architecture rejected 98.7% of potentially dangerous leads.

Framework and Perspectives: Closed Loop and Competition in the Fall

Initially presented in February 2025 on Gemini 2.0 with gaps in verification and review, the system has evolved into a research partner equipped with modules for controlling digital claims. In biology, its integration has included expert feedback. Researchers see this as a step toward closed-loop agents that improve through experimental feedback and could accelerate research, while noting, as Samuel Schmidgall points out, that a long way remains to be traveled to deal with physical realities. Automated research is attracting attention, and OpenAI plans to present an agent operating at least at the level of an intern this fall. A debate persists regarding the ability of LLM-based approaches to genuinely discover new knowledge rather than merely extracting patterns from their training data.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.