Brief IA

Alibaba's Qwen 3.8 Competes with Claude Opus 4.8

🤖 Models & LLM·Tom Levy·

Alibaba's Qwen 3.8 Competes with Claude Opus 4.8

Alibaba's Qwen 3.8 Competes with Claude Opus 4.8
Key Takeaways
1Alibaba's Qwen3.8 Max achieves a score of 56 on the Artificial Intelligence Analysis Index, marking a 10-point improvement.
2Despite this advancement, Kimi K3 still outperforms Qwen3.8 Max and Claude Opus 4.8 in terms of performance.
3Kimi K3 delivers superior performance while reducing costs by 25% compared to its competitors.
💡Why it mattersThe improvement of Qwen3.8 Max demonstrates Alibaba's increasing competitiveness, but Kimi K3 remains the leader in efficiency and cost.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Qwen3.8 Max from Alibaba Competes with Claude Opus 4.8

Qwen3.8 Max scores 56 on the Artificial Intelligence Analysis Index, an increase of 10 points compared to Qwen3.7 Max (46). According to Artificial Analysis, this places it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which is also 25% cheaper.

On the GDPval-AA, a benchmark for work-related tasks, Qwen rises from 468 Elo points to 1,739, surpassing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The issue lies in how it achieves this score. Qwen3.8 Max requires 64 steps per task instead of 14, and the input tokens have increased by 15 times as the test returns the complete conversation history to the model at each step.

The model operates more thoroughly but is slower and more expensive. Alibaba's value for money takes a hit despite the decrease in token prices (input has dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double that of Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 stands at $0.57.

There are also regressions compared to the previous version. The AA-LCR has dropped by 2 points, a test that checks if a model can correctly gather information from very long texts. The AA-Omniscience has decreased by 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it does not know. The accuracy rate remains around 31%, but the hallucination rate has surged from 23% to 40%. Qwen3.8 Max guesses much more often instead of stating that it does not know.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.