⚡
Brief IA
›

LLM Agents' Creativity: GPT-5 Outperforms Embeddings

🔬 Research·Tom Levy·

LLM Agents' Creativity: GPT-5 Outperforms Embeddings

LLM Agents' Creativity: GPT-5 Outperforms Embeddings
⚡
Key Takeaways
1Researchers assess the creativity of LLM agents on MLE-bench by combining P- and H-creativity, impact, and feasibility.
2Evaluation by GPT-5 aligns best with human judges, outperforming embedding-based approaches.
3All agents see their impact increase but their P-creativity decline, quickly shifting from exploration to exploitation.
💡Why it matters — This protocol sheds light on the current limitations of LLM agents compared to human creativity on machine learning research tasks.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Researchers have tested the creativity of LLM agents on MLE-bench, using hundreds of Kaggle solutions as references. It shows that evaluation by GPT-5 aligns best with human annotators and that agents quickly shift from exploration to exploitation. The protocol details how P- and H-creativity, impact, and feasibility are quantified.

An LLM Judge Significantly Aligns with Humans

Three human annotators rated 300 episodes for P-creativity, serving as a reference for automated methods. The Spearman correlations between these methods and human judgments are all significant (p < 0.001). The LLM-as-a-judge approach with GPT-5 achieves the best agreement, ahead of measures based on embeddings. Semantic distance remains effective, but the discrepancies between models used as judges highlight a need for better reasoning capabilities to evaluate novelty. For H-creativity, the size of human corpora imposes a two-step strategy: retrieve 5 nearest neighbors by semantic distance, then solicit the LLM's judgment on these 5 references.

Agents That Improve Impact but Lose P-creativity

Across successive episodes, all agents increase their impact. AIDE with GPT-5 progresses steadily, while AIRA-MCTS with Qwen starts higher but reaches a plateau. P-creativity generally declines over episodes, with systematically lower levels for AIRA-MCTS with Qwen. The tracking begins at episode 1, with episode 0 serving as a reference. As the test time increases, agents shift from exploring new avenues to refining a chosen path. This transition also exists among humans, but agents reach it faster and see their P-creativity decline more quickly.

How Creativity is Defined and Decomposed

Creativity is defined as the production of ideas or products that are both original and useful. Foundational work links it to the exploration of structured conceptual spaces and the ability to target a relevant portion of a vast problem-solving space. Originality is divided into P-Creativity, relative to a program's memory and history, and H-Creativity, appreciated against the corpus of human knowledge. Utility is broken down into impact and feasibility. Together, these form the analytical foundation: P-Creativity, H-Creativity, impact, and feasibility.

Operational Metrics for P-, H-creativity, Impact, and Feasibility

P-Creativity is rated by GPT-5 on a scale from 0 to 4, anchored between the poles of "Routine" and "Transformational." H-Creativity relies on proximity retrieval by semantic distance, followed by a judgment from GPT-5 comparing it to corpora of 877 to 3,747 notebooks per task. Impact is defined by a normalized score reflecting the distance between a baseline and the best human performance. Feasibility is not an explicit score: an episode is only retained if its code executes correctly. The notion of an episode corresponds to a set of steps leading to a successful submission.

A Concrete Case: Cassava, Between New Ideas and Repetition

In the classification of cassava leaf disease images, five consecutive episodes of an AIDE run powered by GPT-5 were observed. The agent starts with a ridge classifier at episode 0, which is conventionally assigned a P-creativity of 4 due to a lack of history. H-creativity and impact peak when it exploits LightGBM with manually designed features, while human competitors would have favored CNNs from the outset. In this context, LightGBM stands out as innovative against 3,747 human approaches. After episode 4, the agent reiterates the same path, favoring exploitation, and finishes far from a medal.

An Evaluation at the Scale of MLE-bench and Human Trajectories

Ten MLE-bench tasks covering image, NLP, and tabular data were selected for their dense human corpora, ranging from 877 to 3,747 public notebooks. Two agents are considered: AIDE, based on a greedy tree search, and AIRA-Dojo, which integrates additional strategies and operators into the former, using GPT-5 and Qwen3-32B as underlying models. Eight executions are performed for each agent and each model-task combination, with a limit of 8 hours and a maximum of 10 episodes. The interest of MLE-bench is to provide human trajectories, allowing for the tracking of impact and P-creativity evolution throughout competitions and directly comparing these dynamics to those of the agents.

The Study Analyzes the Gap Between Announcements and Results of LLM Agents

The work, published at COLM 2026 and co-authored with researchers from the University of Michigan, addresses the gap between breakthrough announcements and the actual performance of LLM agents on machine learning research challenges. While projects like AlphaEvolve or Kosmos are cited and OpenAI claims to have solved the Navier–Stokes equations, the best humans maintain an advantage on concrete problems. One hypothesis posits the role of the framework and scaffolding provided to agents. The authors propose creativity as a key analytical lens and raise the question of attributing performance gaps to the structuring of creative research. ML tasks provide the suitable ground to study, with LLM support, how frameworks guide the emergence of new and useful ideas and whether creativity metrics can explain why certain frameworks or models outperform others.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.