Brief IA

Artificial Analysis Revises Its Index, Astra Advances

🤖 Models & LLM·Tom Levy·

Artificial Analysis Revises Its Index, Astra Advances

Artificial Analysis Revises Its Index, Astra Advances
Key Takeaways
1Artificial Analysis releases the Intelligence Index 4.2 with two new benchmarks and increased weighting for private data
2GPT-6 Astra gains four points and reaches second place behind Claude Fable 5.1
3Version 5 has been in preparation for eight months and will be rolled out in phases
💡Why it mattersThis revision alters the hierarchy of AI models and adjusts the evaluation methodology, impacting performance comparisons.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Artificial Analysis has updated its Intelligence Index to version 4.2, featuring new benchmarks, increased weighting for private data, and scoring corrections. GPT-6 Astra gains four points and reaches second place, behind Claude Fable 5.1. The publisher also indicates that a version 5, in development for eight months, will be rolled out in phases.

Increased Weighting and New Tests Modify the Index

The Intelligence Index now includes two new benchmarks: AA-Briefcase, designed to measure real-world knowledge work, and GDP.pdf from Surge AI, used for evaluating PDF document analysis. The GPQA-Diamond benchmark has been removed, as it has been solved by the evaluated models. Private test data now accounts for 40% of the overall score, aiming to make the evaluation more demanding. Artificial Analysis has also corrected scoring errors across several benchmarks and modified its scoring methods to enhance result stability.

Reshuffled Rankings: Claude Fable 5.1 Leads, Astra Follows

With version 4.2, GPT-6 Astra advances by four points compared to its predecessor. Anthropic's Claude Fable 5.1 retains the top position, followed by GPT-6 Astra in second place and Meta in third. Artificial Analysis notes that GPT-6 Astra uses fewer tokens per task than other leading models. Anthropic, OpenAI, Meta, and Zhipu AI collectively hold the top spot in terms of cost-performance ratio.

Evaluation Gaps and V5 Timeline

Prior to this update, other analyses had already positioned GPT-6 Astra in first place. Epoch AI had ranked it at the top among 267 models with a score of 169 points across more than 50 benchmarks, and ARC-AGI-3 had observed improvements ranging from significant to very significant depending on the systems. Artificial Analysis initially assigned Astra a score similar to that of the previous version, which raised questions. The publisher clarifies that it delayed updates to maintain score stability during major launches but deemed an interim update necessary due to the rapid evolution of the top rankings. Version 5, developed over the past eight months, will be introduced in phases.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.