Nvidia: 100% on ARC-AGI-3 Thanks to a Supervisor Harness

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Researchers at Nvidia achieved 100% on the ARC-AGI-3 test with Claude Opus 5 using a harness tailored for memory and supported by a supervisor. OpenAI claims to have tripled its scores by adjusting two harness parameters, though it did not reach 100%. Other results, including those from Databricks, also highlight the role of the harness in costs. Nvidia, meanwhile, advocates for open building blocks to create agents and argues for an open agent stack.
Databricks warns: a poor harness can double costs
In July, Databricks published findings indicating that the harness, more than the model, significantly impacts AI costs. Its CEO, Ali Ghodsi, warns that by keeping the same model but changing the harness, one could end up with much higher costs if a poor harness is chosen, potentially doubling the bill. These signals add to the notion that the model is not the sole determinant of agent performance. Nvidia emphasizes that open harnesses give users more control, similar to open models. Adel El Hallack argues that open harnesses provide more levers to increase accuracy and connects this approach to the training slowdowns at OpenAI due to security breaches. He advocates for an open agent stack, with control over the harness, infrastructure, and execution, which he deems necessary to advance the ecosystem securely. Nvidia is already offering open components to build harnesses under the Nemo brand, some of which are commercial and a large part freely available.
Claude Opus 5 achieves 100% on ARC-AGI-3 with a supervisor
Researchers achieved 100% with Claude Opus 5 on ARC-AGI-3 by relying on a custom harness optimized for memory and featuring a supervisory component. ARC-AGI-3 includes 2D games without instructions, where the model must understand how to play and win; achieving 100% means winning as well as a human. Without this harness, Opus 5 peaked at 30%, the best score among the tested models. Researchers believe a supervisory component is necessary to guide the agent when it gets stuck. Adel El Hallack emphasizes the value of adding a supervisory agent above the main agent, acting like a CEO to reframe it when it deviates, heads toward a dead end, or re-explores a path. While the idea of a supervisor is not new, most agent users today stick to monolithic harnesses (e.g., Claude Code, Codex, or Hermes). For their trials, Nvidia researchers designed an enhanced harness, the Agent Variation Operators (AVO). AVO is not a product; Nvidia is directing developers toward its Nemo components, a portion of which is freely accessible.
OpenAI tripled its scores by adjusting two harness parameters
OpenAI reported performance below 10% on ARC-AGI-3 for its models, a benchmark that greatly frustrated the company. Last month, the firm conducted its own experiments and demonstrated that by adjusting just two harness parameters, its model scores tripled. However, none reached 100%.
Why the harness weighs more than the model in the long run
Nvidia released research on Friday that places the harness at the forefront of agent systems for long-term tasks. A harness combines tools, memory, and rules, transforming a model into an agent by managing memory, context, and feedback. Adel El Hallack reminds us that many mistakenly equate an agent with a simple model API, whereas it is a broader set: model, harness, tooling, execution, as well as associated skills and libraries. Long-term tasks involve a chain of many decisions, sometimes over several days, where a simple prompt exchange is insufficient; achieving success without drift is considered the Holy Grail of agent research. In April, Microsoft evaluated 19 LLMs on document editing and noted that all, including the most advanced, filled files with errors. Autonomous agents have also been caught deleting entire files or databases and engaging in problematic behaviors, from collusion to hacking. These observations support the idea that, for these uses, the harness often matters more than the model alone.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.