AI Agents: Two Articles Rejected in Open Research

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
As AI players discuss soon-to-be self-improving systems and language models that already write code, generate data, and optimize chips, a study coordinated by Princeton arrives at an inconvenient time. Tested on novel questions drawn from submissions to NeurIPS 2026, agents produced two rejected papers despite six days, $3,000 in API credits, and computational resources. Researchers find these systems proficient in engineering but lacking in judgment and creativity, two ingredients anticipated for recursive improvement.
Warnings Against Self-Improvement Ambitions
The study's results could challenge claims of imminent recursive improvement. In June, Anthropic published a blog post titled "When AI Builds Itself," describing advances towards models capable of accelerating their own development, and in July, OpenAI announced that GPT-5.6 Sol had contributed to post-training a smaller model, reducing researchers' workload. Jack Clark, co-founder of Anthropic, notes that these observations align with what the company has found in attempts to automate AI safety research. He believes current systems lack intuitive creativity and remain mechanical in their reasoning, which he describes as a "downward signal on short-term recursive improvement timelines." Meanwhile, AI companies are seeking to develop systems that can accelerate their own progress; OpenAI explicitly aims for an automated AI researcher, and Anthropic considers self-improvement a key objective. One question remains: is open research essential for recursive improvement, or could progress on more limited tasks suffice?
Researchers Test a "Shadow Evaluation" on NeurIPS 2026
To assess unbounded skills, the multi-institution team led by Peter Kirgis and Sayash Kapoor (Princeton) designed a "shadow evaluation": having an agent respond to a question extracted from a high-quality unpublished paper. Anthropic's Claude Opus 4.8, orchestrated via the open-source tool OpenClaw, was tasked with two questions from submissions to NeurIPS 2026: controlling the "personas" of an LLM by modifying weights, and detecting reliability loss in a model operating on spreadsheet data. Since the source papers were not public, no answers could be retrieved from training or the web. The agents had six days, $3,000 in API credits, a GPU budget, dedicated virtual machines, and access to the open web to produce a paper at the level of a major conference. The authors of the original papers evaluated these outputs as genuine submissions.
Two Papers Rejected Despite Strong Technical Execution
The two papers produced by the agents were rejected. Human evaluators noted a good command of engineering: literature review, conducting numerous experiments, and compiling results. The agents can thus solve technical tasks necessary for research, but according to the team, they lack the judgment and creativity required to reach the level of top conferences. Sayash Kapoor describes them as "unequivocally poor" in research conduct: unusual experiments, sometimes on small synthetic datasets, difficult-to-understand writing, and a lack of innovative contributions, far from the expected standard. The agents explored little, committed too quickly to unpromising paths, designed ambitious hypotheses similar to those of humans, then rejected them based on very limited data, without a real ability to backtrack or start over. They did not incorporate feedback from sub-agents or external evaluation tools, restricted their claims instead of revising their methods, poorly managed tokens, computation, and time, and did not adhere to guidelines on time allocation or paper length. However, researchers noted the absence of "reward manipulation": while some sub-agents hallucinated or distorted results, the orchestrating agent detected these errors.
Kapoor Links Observed Gap to Training Methods
Sayash Kapoor argues that models excel primarily where reinforcement learning can be applied and where success is automatically verifiable. Designing such environments becomes more challenging when the task is open, which could explain correct performance in engineering but poor performance in open research.
Limited Scope of the Study and New Trial Underway with Mythos
The study focused on only two papers, and the original authors knew they were evaluating AI-generated texts, which may have influenced their judgments. Researchers had considerable latitude in design and execution, with possible biases; open research evaluation sacrifices some objectivity for a richer test than a benchmark. The team is now conducting the experiment with Mythos, presented as Anthropic's most advanced model, launched in April, then subjected to security restrictions imposed by the Trump administration and reserved for approved organizations. Anthropic did not respond to a request for comment.
Industry Promises and Possible Trajectories Discussed by Najoung Kim
As the industry projects rapid self-improvement and highlights existing capabilities of LLMs such as coding, data generation, and hardware optimization, many works primarily measure narrow and verifiable tasks. However, open research involves choosing hypotheses, determining relevant evidence, and knowing how to restart. Najoung Kim believes targeted investments could enable progress while envisioning a bifurcated trajectory: rapid on notable tasks and slower on open research. Sayash Kapoor reminds that major advances like the invention of transformers or new architectures required creative leaps, while others believe existing levers—training faster and improving scores—might suffice. The observed gap suggests that some timelines for automating research may be premature, and the study underscores that achieving recursive improvement could take longer than anticipated.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.