Qwen 3.5 and GLM 5: Chinese AI Giants on the Rise
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A Month Rich in AI Innovations in China
The past month has been particularly dynamic for the artificial intelligence sector in China, with the launch of several flagship models. Among them, Qwen 3.5, MiniMax 2.5, and GLM 5 have captured the attention of experts and the public. Anticipation is also high for the upcoming release of DeepSeek V4, fueled by persistent rumors. Outside of the cutting-edge models, this month has been relatively quiet in terms of new niche modalities and model sizes.
Introduction of Relative Adoption Metrics
To track these new releases, a new measurement tool has been introduced: Relative Adoption Metrics (RAM). This tool allows for the normalization of model downloads relative to their peers in the same size category. It has already proven extremely useful, highlighting underestimated models like GPT-OSS, which has literally exploded in terms of downloads. Since Llama 3.1, it has become the most popular American open-weight model. A RAM score above 1 indicates that a model is on track to become one of the ten most downloaded models of all time in its size category.
Qwen 3.5: A Significant Advancement
The Qwen3.5-397B-A17B model from Qwen has been highly anticipated. It is available in several sizes, ranging from 0.8B to 27B for dense models and from 35B-A3B to 397B-A17B for MoE models. All these models are multi-modal, utilize default reasoning, and are based on the Qwen-Next architecture with GDN layers. After testing conducted in recent days, these models show a marked improvement over the previous version. They are particularly well-suited for a wide range of tasks, with notable enhancements in style and the ability to follow instructions. Additionally, they have improved for multilingual tasks, covering a greater number of languages. However, smaller models still tend to overthink, although this feature can be disabled.
Remarkable Performances from StepFun and Z.ai
StepFun has also made waves with its Step-3.5-Flash model, a 196B-A11B MoE model that demonstrates strong performance across the board, notably outperforming models several times larger in mathematical benchmarks. Meanwhile, Z.ai launched GLM-5, a 744B-A40B model that generated such demand that the Zhipu team had to raise the prices of their coding plan. This model is accompanied by a detailed technical report.
MiniMax 2.5 and Other Notable Models
Despite its relatively small size, MiniMax-M2.5 from MiniMaxAI has managed to compete with models such as GLM-5 and Kimi K2.5, quickly becoming a favorite in the community. OpenThinker-Agent-v1 by open-thoughts marks the entry of OpenThinkers into the realm of agentic reasoning, with an initial release that includes SFT and RL data, as well as a "lite" version of terminal-based tasks to evaluate smaller models.
DeepSeek V3.2 and Other Significant Contributions
DeepSeek V3.2, along with other recent models from the same series, has underperformed compared to the early releases of DeepSeek in 2025, which has been a disappointment for some observers. Among other notable models, Tri-21B-Think by Trillion Labs, a 21B reasoning model, supports English, Korean, and Japanese. MiniCPM-SALA by openbmb, on the other hand, is an 8B model in English and Chinese with sparse attention, supporting a context window of 1M. These innovations underscore the growing importance of open-weight models in the AI ecosystem, with major implications for the future of technology.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.