Z.ai launches GLM-5.3-Flash: cost cut, no Nvidia GPU required

Z.ai is commercializing GLM-5.3-Flash, a model with 320 billion parameters and a context window of 1 million tokens, touted as operable on Chinese chips with efficiency comparable to Nvidia GPUs. Third-party measurements place it near GLM-5.3 at $0.09 per task, and a report mentions up to 100 trillion tokens served daily.
Claimed Capability and Bet on Independence from CUDA
SemiAnalysis reports that the infrastructure associated with the model has served 100 trillion tokens per day. The model operates entirely on Chinese AI chips, and Z.ai claims efficiency comparable to that of current Nvidia GPUs, both in hardware performance and cost per token. SemiAnalysis sees this as a new test of the CUDA gap following recent results around new OpenAI chips. CUDA, Nvidia's programming layer between AI software and GPUs, has been developed for nearly 20 years, and most frameworks are optimized for it, making migration to other chips costly in terms of rewriting operations, memory adjustments, and resolving bottlenecks. To circumvent these constraints, Z.ai built a service based on SGLang and segmented processing into independent steps; the team claims to have tripled throughput compared to its initial implementation on the same hardware, with the help of an agent derived from GLM-5.3 for optimization.
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Measured Performance and Cost per Task at $0.09
Artificial Analysis rates GLM-5.3-Flash at 57 points on the Intelligence Index with maximum reasoning effort, three points behind GLM-5.3 at 60, and on par with GPT-5.6 Terra and Muse Spark 1.2. The cost per task on this index reaches $0.09, about 7.5 times less than GLM-5.3; Artificial Analysis places it on the Pareto frontier between intelligence and cost. On GDPval-AA v2, the Elo of approximately 1770 equals GLM-5.3 and Grok 4.6, trailing Claude Opus 5. However, measurements indicate that about 90% of output tokens are allocated to reasoning.
Specifications: 320 Billion Parameters, 18 Billion Active
GLM-5.3-Flash incorporates a total of 320 billion parameters as well as a context window that can reach 1 million tokens. According to Z.ai, this is the first native multimodal model in the GLM-5 series, with a total of 320 billion parameters, of which 18 billion are activated. The model is released under the MIT license, and its weights are available for download on Hugging Face.
API Pricing and Anonymous Testing Before Launch
On Z.ai's API, the announced rates are $0.15 per million input tokens and $0.50 per million output tokens, for a price slightly above ten percent of GLM-5.3. Before the official announcement, the model was anonymously evaluated under the name "ox-alpha" on OpenCode and OpenRouter, where it became the most popular model of the week. Z.ai indicates that all of this traffic was processed on Chinese AI chips.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.