Anthropic Claude Opus 4.7: Breakthrough in Coding, Cybersecurity Limited
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Claude Opus 4.7: A Breakthrough in Coding
Anthropic recently unveiled Claude Opus 4.7, a model that represents a significant advancement in the field of autonomous coding. This model achieved an impressive score of 64.3% on the SWE-bench Pro benchmark, surpassing its predecessor, Opus 4.6, which scored 53.4%. In comparison, OpenAI's GPT-5.4 scored 57.7%, placing Opus 4.7 in a favorable position.
However, the Claude Mythos Preview model remains at the top with a score of 77.8%. Anthropic emphasizes that Opus 4.7 follows instructions with increased accuracy compared to previous versions. This means that prompts designed for earlier models may now yield unexpected results, as Opus 4.7 interprets instructions more literally.
Significant Visual Improvements
Opus 4.7 introduces a tripled image resolution, capable of processing images up to 2,576 pixels on the long edge, or approximately 3.75 megapixels. This enhancement is particularly beneficial for computing agents that need to analyze complex screenshots and extract data from detailed diagrams. On the Document Reasoning benchmark, Opus 4.7 achieved an accuracy of 80.6%, compared to 57.1% for Opus 4.6.
This improvement is not just a simple API parameter but a fundamental change at the model level, allowing for automated processing of images at a higher resolution, although this consumes more tokens as a result.
Reduction of Cyber Capabilities
A notable aspect of this release is the deliberate reduction of certain cyber capabilities during training. Anthropic has implemented new safeguards to automatically detect and block requests suggesting prohibited or high-risk cyber usage. This strategy is part of the Glasswing project, where Anthropic has discussed the risks and benefits of AI models for cybersecurity.
Opus 4.7 is the first model to test this approach, and interested security researchers can sign up for a new cyber verification program to use the model in penetration testing or red-teaming.
Managing Hallucinations
According to Anthropic's system map, Opus 4.7 shows improvements in managing factual and input hallucinations. Factual hallucinations concern incorrect claims about the world, while input hallucinations occur when the model acts as if it has access to a non-existent tool or attachment.
Opus 4.7 performs better or at the same level as Opus 4.6 on several benchmarks for factual hallucinations, although it still falls short of Mythos Preview for obscure facts. For input hallucinations, Opus 4.7 achieves the lowest hallucination rate of all tested models when users request a tool that is not available.
Alignment and Costs
Overall, Anthropic describes the security profile of Opus 4.7 as similar to that of Opus 4.6, with low rates of deception, sycophancy, and cooperation with abuse. The model is more resistant to prompt injection attacks. However, a persistent issue with earlier Claude models is the refusal to assist in legitimate AI security research. Opus 4.7 still refuses to assist in 33% of simulated security research tasks, although this represents a significant decrease from 88% with Opus 4.6.
Token prices remain at $5 per million input tokens and $25 per million output tokens. However, Opus 4.7 uses a new tokenizer that can map the same text to up to 1.35 times more tokens. The model also generates more output tokens at higher effort levels, meaning that the cost per request can increase significantly even if token prices remain unchanged.
A new effort level called "xhigh" sits between "high" and "max." Claude Code also receives a new command "/ultrareview" for dedicated code reviews and an expanded "Auto Mode" for Max users, where Claude makes decisions autonomously. Opus 4.7 is available via the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
For more details and guidance, users can refer to the migration guide for Opus 4.7.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.