Brief IA

Google Gemma 4: Spectacular 72% Reduction and 4-Bit Bug Fixed

🔬 Research·Tom Levy·

Google Gemma 4: Spectacular 72% Reduction and 4-Bit Bug Fixed

Google Gemma 4: Spectacular 72% Reduction and 4-Bit Bug Fixed
Key Takeaways
1Google has successfully reduced the Gemma 4 model by 72%, making it compatible with 15 GB of memory while maintaining a speed of 193 tokens per second.
2Unsloth has fixed a scaling disagreement bug in GGUF conversion, improving the efficiency of 4-bit quantized models.
3The 26B-A4B model of Gemma 4 offers performance close to the original despite 4-bit quantization, defying traditional expectations.
💡Why it mattersThis advancement allows large models to run on consumer-grade configurations, broadening access to advanced AI technology.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google recently unveiled a major advancement with its Gemma 4 model, which has been reduced by 72% while maintaining impressive performance. This model, equipped with 26 billion parameters, now operates efficiently with just 15 GB of memory, producing 193 tokens per second on consumer GPUs. This technical feat allows for the execution of large models on laptops and gaming setups, rather than in data centers.

The innovation lies in the new QAT checkpoints of Gemma 4, which enable quantization during training. This makes the model robust against rounding in 4 bits, a task that, according to traditional quantization laws, should result in a significant loss of performance. However, this is not the case here.

Unsloth played a crucial role in fixing a subtle scale disagreement bug in the GGUF conversion. This bug, if left unresolved, would have negated most of the benefits when converting to llama.cpp formats. Thanks to this fix, 4-bit quantized models maintain remarkable efficiency.

The article also details the performance and memory of the various Gemma 4 variants, focusing on the 26B-A4B mixture of experts model. It compares naive and dynamic conversions and provides practical steps for running the model with llama.cpp, as well as other deployment options like API servers, Ollama/LM Studio, and Unsloth Studio.

In conclusion, while the 4-bit model remains limited by its nature, the usual trade-off between quality and speed seems to be collapsing. The 26B-A4B model offers an experience comparable to that of large models on consumer GPUs, opening new avenues for accessible AI.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.