Brief IA

OpenAI: 30% of SWE-Bench Pro Tasks Found Deficient

💻 Code & Dev·Tom Levy·

OpenAI: 30% of SWE-Bench Pro Tasks Found Deficient

OpenAI: 30% of SWE-Bench Pro Tasks Found Deficient
Key Takeaways
1OpenAI analyzed the SWE-Bench Pro coding test, revealing that 30% of the tasks are failing.
2The company has decided to withdraw its support for this test, which was once widely used to evaluate AIs.
3This finding calls into question the reliability of tools used to assess AI programming skills.
💡Why it mattersOpenAI's questioning of SWE-Bench Pro could lead to a revision of AI evaluation standards, impacting developers and tech companies.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI: 30% of SWE-Bench Pro Tasks Found Defective

OpenAI has withdrawn its support for the AI coding test SWE-Bench Pro after discovering that approximately 30% of its tasks are defective.

The issues arise from the fact that the tasks are drawn from real software projects, making them too strict, too vague, or misleading for AI models. This skews the assessment of what AI can actually achieve.

OpenAI is calling for more reliable benchmarks. Artificial Analysis had already removed the test from its rankings after finding that some models were copying solutions from project commit histories instead of solving the tasks.

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring the programming skills of AI models, and found that about 30% of its tasks were defective. The company is retracting its previous support for this benchmark.

The results of tests like this influence decisions regarding the release of a model, including safety assessments as part of OpenAI's Preparedness Framework. When a test contains errors, it can give a misleading picture of what AI can actually do.

To conduct this evaluation, OpenAI first deployed an automated filtering tool that flagged 286 suspicious tasks. Codex-based AI agents then examined each case in detail before a human researcher made the final decision. This process classified 200 tasks (27.4%) as defective. In a parallel assessment, five experienced software developers reviewed the same cases and reported even more, 249 tasks (34.1%). Human evaluators were stricter than the AI agents, although both parties agreed in 74% of the cases.

A Single Space Character Can Mean Success or Failure

OpenAI categorizes the problems into four groups:

  • Some tests are too strict, rejecting solutions that actually work.
  • Others are too vague, expecting the AI to meet requirements buried in hidden test cases.
  • Some tests are too superficial, allowing incomplete solutions to pass.
  • And some task descriptions simply point in the wrong direction.

One example from the OpenLibrary project: the task description required a single space, but the hidden test expected two. An AI that followed the instructions correctly would fail.

The tasks were extracted from the commit histories of real software projects, originally written for human collaboration and not designed as clean evaluation tasks for AI models. According to OpenAI, tests derived from these projects tend to be too strict because they were built to verify a specific change, not to serve as general requirements.

On the public version of the test with 731 tasks, the best models improved from 23.3% to 80.3% accuracy in just eight months. SWE-Bench Pro was intended to replace the older SWE-bench Verified, which OpenAI had already rejected for similar reasons.

This time, OpenAI is not recommending a specific replacement. The company is simply calling on the industry to build new benchmarks using experienced developers—those who are hard to circumvent, trustworthy, and genuinely meaningful.

By mid-June, the analysis firm Artificial Analysis had already removed SWE-Bench Pro from its Coding Agent Index and replaced it with DeepSWE, a test from Datacurve. The reason: SWE-Bench Pro was manipulable. Some models had copied the correct solution from a project's commit history instead of actually solving the task.

This change reshuffled the rankings. Codex with GPT-5.5 (xhigh) rose from 65 to 76 points and surpassed Claude Code with Opus 4.8 (max) at 73, while Claude Code with Fable 5 (max) took the top spot with 77 points. On SWE-Bench Pro, Codex with GPT-5.5 had only scored 31 points, compared to 64 to 84 on other tests.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.