Claude Opus 4.8 and GPT-5.6 Sol: Rejected Research Papers

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
In an autonomous setting for writing scientific articles, agents equipped with Claude Opus 4.8 and GPT-5.6 Sol produced texts that were rated "Rejected" by the original authors of unpublished NeurIPS papers. According to the study, conducted with Princeton and the UK AI Security Institute, the models can execute the research engineering process but stumble on judgment and creative problem-solving, as well as on abandoning ineffective approaches. These findings contradict the claims made by Anthropic and OpenAI, which present autonomous research as imminent.
The study highlights gaps in judgment and creativity
The study indicates that state-of-the-art models are capable of carrying out the entire research engineering process. However, it points out failures in scientific judgment, creative problem-solving, and the ability to abandon unproductive paths. The generated articles were ultimately rated as "Rejected."
Six days, $3,000, and GPUs for Claude and GPT agents
The AI agents utilized Claude Opus 4.8 and GPT-5.6 Sol. They were provided with six days, $3,000 in API credits, and GPU access to autonomously write AI research papers. The texts produced were evaluated by the original authors of unpublished NeurIPS papers.
With Princeton and the UK AISI, claims are challenged
The study was conducted in collaboration with Princeton and the UK AI Security Institute. Its findings oppose the claims made by Anthropic and OpenAI, which present autonomous AI research as imminent.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.