OpenAI: 30% of SWE-Bench Pro Tasks Found Deficient

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI: 30% of SWE-Bench Pro Tasks Found Defective
OpenAI has withdrawn its support for the AI coding test SWE-Bench Pro after discovering that approximately 30% of its tasks are defective.
The issues arise from the fact that the tasks are drawn from real software projects, making them too strict, too vague, or misleading for AI models. This skews the assessment of what AI can actually achieve.
OpenAI is calling for more reliable benchmarks. Artificial Analysis had already removed the test from its rankings after finding that some models were copying solutions from project commit histories instead of solving the tasks.
OpenAI reviewed SWE-Bench Pro, a widely used test for measuring the programming skills of AI models, and found that about 30% of its tasks were defective. The company is retracting its previous support for this benchmark.
The results of tests like this influence decisions regarding the release of a model, including safety assessments as part of OpenAI's Preparedness Framework. When a test contains errors, it can give a misleading picture of what AI can actually do.
To conduct this evaluation, OpenAI first deployed an automated filtering tool that flagged 286 suspicious tasks. Codex-based AI agents then examined each case in detail before a human researcher made the final decision. This process classified 200 tasks (27.4%) as defective. In a parallel assessment, five experienced software developers reviewed the same cases and reported even more, 249 tasks (34.1%). Human evaluators were stricter than the AI agents, although both parties agreed in 74% of the cases.
A Single Space Character Can Mean Success or Failure
OpenAI categorizes the problems into four groups:
- Some tests are too strict, rejecting solutions that actually work.
- Others are too vague, expecting the AI to meet requirements buried in hidden test cases.
- Some tests are too superficial, allowing incomplete solutions to pass.
- And some task descriptions simply point in the wrong direction.
One example from the OpenLibrary project: the task description required a single space, but the hidden test expected two. An AI that followed the instructions correctly would fail.
The tasks were extracted from the commit histories of real software projects, originally written for human collaboration and not designed as clean evaluation tasks for AI models. According to OpenAI, tests derived from these projects tend to be too strict because they were built to verify a specific change, not to serve as general requirements.
On the public version of the test with 731 tasks, the best models improved from 23.3% to 80.3% accuracy in just eight months. SWE-Bench Pro was intended to replace the older SWE-bench Verified, which OpenAI had already rejected for similar reasons.
This time, OpenAI is not recommending a specific replacement. The company is simply calling on the industry to build new benchmarks using experienced developers—those who are hard to circumvent, trustworthy, and genuinely meaningful.
By mid-June, the analysis firm Artificial Analysis had already removed SWE-Bench Pro from its Coding Agent Index and replaced it with DeepSWE, a test from Datacurve. The reason: SWE-Bench Pro was manipulable. Some models had copied the correct solution from a project's commit history instead of actually solving the task.
This change reshuffled the rankings. Codex with GPT-5.5 (xhigh) rose from 65 to 76 points and surpassed Claude Code with Opus 4.8 (max) at 73, while Claude Code with Fable 5 (max) took the top spot with 77 points. On SWE-Bench Pro, Codex with GPT-5.5 had only scored 31 points, compared to 64 to 84 on other tests.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.