Artificial Analysis Revises Its Index, Astra Advances

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Artificial Analysis has updated its Intelligence Index to version 4.2, featuring new benchmarks, increased weighting for private data, and scoring corrections. GPT-6 Astra gains four points and reaches second place, behind Claude Fable 5.1. The publisher also indicates that a version 5, in development for eight months, will be rolled out in phases.
Increased Weighting and New Tests Modify the Index
The Intelligence Index now includes two new benchmarks: AA-Briefcase, designed to measure real-world knowledge work, and GDP.pdf from Surge AI, used for evaluating PDF document analysis. The GPQA-Diamond benchmark has been removed, as it has been solved by the evaluated models. Private test data now accounts for 40% of the overall score, aiming to make the evaluation more demanding. Artificial Analysis has also corrected scoring errors across several benchmarks and modified its scoring methods to enhance result stability.
Reshuffled Rankings: Claude Fable 5.1 Leads, Astra Follows
With version 4.2, GPT-6 Astra advances by four points compared to its predecessor. Anthropic's Claude Fable 5.1 retains the top position, followed by GPT-6 Astra in second place and Meta in third. Artificial Analysis notes that GPT-6 Astra uses fewer tokens per task than other leading models. Anthropic, OpenAI, Meta, and Zhipu AI collectively hold the top spot in terms of cost-performance ratio.
Evaluation Gaps and V5 Timeline
Prior to this update, other analyses had already positioned GPT-6 Astra in first place. Epoch AI had ranked it at the top among 267 models with a score of 169 points across more than 50 benchmarks, and ARC-AGI-3 had observed improvements ranging from significant to very significant depending on the systems. Artificial Analysis initially assigned Astra a score similar to that of the previous version, which raised questions. The publisher clarifies that it delayed updates to maintain score stability during major launches but deemed an interim update necessary due to the rapid evolution of the top rankings. Version 5, developed over the past eight months, will be introduced in phases.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.