Google Gemma 4: Spectacular 72% Reduction and 4-Bit Bug Fixed

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Google recently unveiled a major advancement with its Gemma 4 model, which has been reduced by 72% while maintaining impressive performance. This model, equipped with 26 billion parameters, now operates efficiently with just 15 GB of memory, producing 193 tokens per second on consumer GPUs. This technical feat allows for the execution of large models on laptops and gaming setups, rather than in data centers.
The innovation lies in the new QAT checkpoints of Gemma 4, which enable quantization during training. This makes the model robust against rounding in 4 bits, a task that, according to traditional quantization laws, should result in a significant loss of performance. However, this is not the case here.
Unsloth played a crucial role in fixing a subtle scale disagreement bug in the GGUF conversion. This bug, if left unresolved, would have negated most of the benefits when converting to llama.cpp formats. Thanks to this fix, 4-bit quantized models maintain remarkable efficiency.
The article also details the performance and memory of the various Gemma 4 variants, focusing on the 26B-A4B mixture of experts model. It compares naive and dynamic conversions and provides practical steps for running the model with llama.cpp, as well as other deployment options like API servers, Ollama/LM Studio, and Unsloth Studio.
In conclusion, while the 4-bit model remains limited by its nature, the usual trade-off between quality and speed seems to be collapsing. The 26B-A4B model offers an experience comparable to that of large models on consumer GPUs, opening new avenues for accessible AI.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.