Alibaba and DeepSeek: Low-Cost Chinese AI

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Alibaba and DeepSeek: Making AI Models More Accessible in China
Alibaba recently unveiled Qwen3.8-Max, its most advanced artificial intelligence model to date. This model stands out with its impressive size of 2.4 trillion parameters. Thanks to a mixture of experts architecture, it activates only a fraction of its parameters for each request, significantly reducing costs and response times. On average, 95 billion parameters are active simultaneously, optimizing efficiency without requiring the full activation of the model.
Meanwhile, DeepSeek has also adopted a similar architecture but on a more modest scale. The V4-Flash model is listed by Artificial Analysis with 284 billion parameters, of which 13 billion are active during inference. In comparison, Moonshot AI's Kimi K3, launched in July, has 2.8 trillion parameters with about 104 billion active.
Competition on Features and Costs
Alibaba's Qwen3.8-Max is capable of processing text, images, and videos, while supporting up to one million tokens of context. Alibaba also claimed that this model successfully completed a software engineering project in just 16 days. In terms of cost, Qwen3.8-Max is priced at $2 per million input tokens and $6 per million output tokens, which is competitive compared to $3 and $15 respectively for the Kimi K3 model.
The size of a model is not the only factor influencing inference costs. Other elements such as architecture, the number of active parameters, token consumption, and the number of calls needed to complete a task also play a crucial role in the total execution cost of a model.
Rankings and Model Performance
After its launch, Qwen3.8-Max quickly topped the Chinese text models on the comparison platform Arena.AI. However, it remained behind several models from Anthropic in the overall ranking. Regarding image analysis and other visual materials, it ranked second, just behind a version of Anthropic Claude Fable 5.
DeepSeek: An Aggressive Pricing Strategy
DeepSeek has chosen a different approach with its V4-Flash model, focusing on lower inference rates than many competing AI systems. According to Artificial Analysis, V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens. The model has a context window of one million tokens and 284 billion parameters, of which 13 billion are active during inference.
A cache rate of $0.003 per million tokens is offered for the Max Effort version of V4-Flash, representing a 98% reduction compared to the standard entry rate. These lower rates are reflected in the benchmark tests from Artificial Analysis, which estimated the average cost of V4-Flash at three cents per test, compared to 86 cents for Kimi K3, $1.86 for OpenAI's GPT-5.6 Sol, and $3.15 for Claude Fable 5 from Anthropic.
The Impact of Costs on Performance
Moonshot AI's Kimi K3 illustrates how announced API prices can differ from the actual cost of executing longer tasks. Artificial Analysis lists this model at $3 per million input tokens and $15 per million output tokens, with cached inputs at $0.30 per million tokens. On the AA-Briefcase benchmark for agentic knowledge work, Kimi K3 averaged $10.57 per task, generating about 120,000 output tokens and using an average of 83 turns per task.
The cost reflects Kimi K3's token pricing, output volume, and the number of interactions with the model. Repeated calls and larger outputs can thus increase the total cost beyond what the main API rate suggests.
The Option of Open Weights
The cost of using models is also influenced by how they are distributed. Alibaba, DeepSeek, and Moonshot AI continue to support open-weight versions alongside hosted API access, providing developers with more flexibility for deployment.
Artificial Analysis lists DeepSeek V4-Flash as an open-weight model under an MIT license, with weights available via Hugging Face. Kimi K3 is also available as an open-weight model under Moonshot AI's own license.
Open weights allow developers to run models on their own infrastructure or through third-party providers, rather than relying solely on a developer-hosted inference service. Deployment costs depend on the hardware and infrastructure used, but access to the model is not tied to a single hosted API.
This approach contrasts with the main models offered by OpenAI, Anthropic, and Google, which generally keep their model weights closed.
Lian Jye Su, chief analyst at Omdia, emphasizes that the choice of model for many commercial workloads does not solely depend on access to the highest-performing system. According to Su, “many commercial workflows do not need the best model in the industry. They need models that are good enough, affordable, transparent, and accessible, and open-weight models help meet that demand.”
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.