⚡
Brief IA
›

Google Details RSI for Better Generalizing Agents

🔬 Research·Tom Levy·

Google Details RSI for Better Generalizing Agents

Google Details RSI for Better Generalizing Agents
⚡
Key Takeaways
1RRSI, developed by Google Cloud AI Research and universities, guides the self-publishing of agent harnesses to prevent test memorization
2The team reports up to a 4.7-point gain on unseen benchmarks and about 30% fewer tokens during execution
3The discovered mechanisms transfer to weaker models, and the code is available
💡Why it matters — RRSI aims to enhance the robustness and generalization of AI agents while reducing computational costs.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Researchers from Google Cloud AI Research and several universities have introduced RRSI, a self-improvement procedure for agents that avoids test memorization. The team reports gains on unprecedented benchmarks, a reduction of about 30% in execution tokens, and improvements transferable to weaker models, with supporting code published.

Gains on 8 Benchmarks and 30% Fewer Tokens

The research team tested RRSI on eight benchmarks encompassing programming, agentic office tasks, and engineering design, using the Claude Opus 4.8 model without modifications. Compared to an unmodified baseline harness and four recent methods, RRSI achieves gains of up to 14.1 points on training tasks and up to 4.7 points on five unseen benchmarks. The team also reports approximately 30% fewer tokens during execution compared to the unregularized variant. Overall performance never fell below the baseline on unseen benchmarks, and every out-of-training task was improved, peaking at 4.7 points on JobBench. An optimized programming harness with Gemini 3.5 Flash also boosted the accuracy of the weaker Gemini 3.1 Flash Lite by 11.2 to 14.6 points without any other modifications. The mechanisms uncovered do not depend on the model's capacity to discover them.

Agents Lose Generalization After Repeated Self-Improvement Cycles

When agents self-optimize in a loop on a limited set of test tasks, they eventually memorize these exercises. Their training scores rise, while gains on new tasks diminish or disappear. The researchers describe several mechanisms at play: the search anchors patterns specific to a benchmark, it may favor high-performing candidates by chance, and it sometimes adds complexity that artificially inflates scores. Many previous approaches primarily enhance training, with little transfer to unseen benchmarks. RRSI is proposed to counter these pitfalls and improve scores across these three dimensions.

RRSI Frames Harness Rewrites Without Locking Them In

Agents rely on a harness that orchestrates prompts, tools, memory, flow, and logic, including verifying which file to open, how to recover from an error, and how to present results. Recent methods assign a language model successive rewrites of this harness based on feedback, in a form of self-directed recursive improvement. RRSI operates at both ends of this cycle. On one side, a budget limits the number of independent modifications a proposal can group, and this budget decreases over time, shifting from broad overhauls to small changes explicitly linked to outcomes. The system logs attempts to avoid re-launching failed paths and explores new areas of the harness in case of stagnation. On the other side, the promotion of changes to permanent elements is filtered by a critic that excludes any hardcoding of task names, solutions, or benchmark-specific tricks. Higher computational costs are only accepted with measurable gains, and obsolete components are removed.

Scope of the Study, Conclusion, and Code Availability

The trials are limited to harnesses based on models with fixed parameters, without addressing situations where weights are modified. They conclude that self-improvement reliably enhances agents' capabilities when repeated feedback leads to lasting changes in the harness. The GitHub repository provides access to the associated code.

Context: Difficult Generalization and Related Approaches

Trials on ARC-AGI-3 have illustrated the limits of generalization of manually designed harnesses: with a custom harness, Opus 4.6 achieved 97.1% in a familiar environment versus 0% in an unfamiliar environment. Other work also aims to contain costs and improve agents' robustness. Nvidia has presented SoL-Pi, where a research agent automatically reconstructs the harness of programming agents, reducing token usage by up to 49% without a notable drop in performance. Shortly before, Google teams had "dreamed" agents of their past research to refine their strategy while keeping the model unchanged. In this landscape, Google Cloud AI Research and universities introduce RRSI as a way to limit test memorization while reducing costs.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.