Meta Revolutionizes Local AI with Muse Glimmer on Consumer GPUs

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Meta Muse Glimmer: A Leap Forward for Local AI
Meta recently unveiled Muse Glimmer, an AI agent model designed to operate efficiently on consumer-grade GPUs. Released under an Apache 2.0 license, this 30 billion parameter model is now accessible via the Hugging Face platform. Meta has highlighted the use of Muse Glimmer for various applications, such as local coding, function calling, and evaluating large language models (LLMs) as judges. This initiative aims to overcome the limitations of cloud-hosted models, which require a constant network connection and centralized infrastructure. By offering Muse Glimmer, Meta provides a solution for workloads that require on-device processing, including personal agents capable of accessing private information like calendars, messages, and files.
Top Performance in Agent Benchmarks
Performance tests conducted by Meta have placed Muse Glimmer at the top of several general agent benchmarks. For instance, on the MCP Atlas benchmark, Muse Glimmer achieved an impressive score of 75.5, surpassing Gemma4-31B, which scored 54.2, and Qwen3.6-27B with 62.5. In the DeepSearch QA test, Muse Glimmer also dominated with a score of 74.6, compared to 61.7 for Gemma4-31B and 71.1 for Qwen3.6-27B. These benchmarks assess the agents' ability to execute complex tasks and handle multi-turn requests. On τ²-Banking, Muse Glimmer recorded 23.5, while Gemma4-31B and Qwen3.6-27B scored 15.1 and 16.7, respectively. Additionally, in the WildClawBench benchmark, Muse Glimmer scored 47.6, outperforming its competitors.
However, some benchmarks favored Qwen3.6-27B, particularly on GDPval-AA, where it scored 1,141 points compared to 953 for Muse Glimmer. On SkillsBench, Qwen3.6-27B also took the lead with a score of 46.6, while Muse Glimmer scored 44.3. The OSWorld-Verified benchmark showed the largest gap, with Qwen3.6-27B at 75.6, Muse Glimmer at 65.9, and Gemma4-31B at 58.5. These results demonstrate the competitiveness of Muse Glimmer, although they do not necessarily reflect its behavior once integrated into enterprise systems.
Comparison of Coding Capabilities
In the coding domain, Muse Glimmer has shown remarkable performance. On the SWE-Bench Pro benchmark, it scored 51.2, surpassing Gemma4-31B at 36.9 and Qwen3.6-27B at 50.2. On SciCode, Muse Glimmer slightly edged out Gemma4-31B with a score of 43.6 versus 43.4, while Qwen3.6-27B scored 39.8. However, Qwen3.6-27B dominated other assessments, including SWE-Bench Verified with 77.2 compared to 76.0 for Muse Glimmer, and TerminalBench 2.1 with 60.7, where Muse Glimmer scored 51.7.
A local coding agent must be capable of more than just generating code. It should be able to interact with repositories, terminals, and testing environments. Muse Glimmer supports agent orchestration models like OpenClaw, and its documentation develops custom structures for developers. Companies need to define the commands and repositories accessible to the agent to evaluate its performance. The model includes mechanisms to handle failed tool calls, requiring checks on repeated attempts to avoid unwanted modifications to the source code.
Multimodal Performance: Text and Images
Muse Glimmer is also designed to handle multimodal data, integrating text and images through a dedicated perception encoder. This design allows agents to interpret screenshots, graphics, and documents within conversations. In the Charxiv Reasoning benchmark, Muse Glimmer scored 78.8, surpassing Gemma4-31B at 77.7 and Qwen3.6-27B at 78.4. However, Qwen3.6-27B excelled in other tests like ScreenSpot Pro with 76.1, where Muse Glimmer scored 75.4.
These results are crucial for teams considering the use of agents capable of interacting with visual interfaces. While Muse Glimmer shows promising capabilities, local tests must still evaluate permissions, display layouts, and potential errors from connected tools.
Security: A Risk Assessment
Meta has also evaluated Muse Glimmer on security criteria, including tests from CI Memories and Siren AgentDojo. On CI Memories, Muse Glimmer recorded a violation rate of 26.4 and a coverage of 64.8, while Gemma4-31B showed a violation rate of 12.1 and a coverage of 53.0. Qwen3.6-27B displayed a violation rate of 53.4 with a coverage of 66.9.
In Siren AgentDojo, Muse Glimmer achieved an attack success rate of 28.4 and a utility score of 94.2. Gemma4-31B scored 25.6 for the attack success rate and 90.8 for utility, while Qwen3.6-27B recorded 40.3 and 92.7, respectively.
General Reasoning: A Nuanced Comparison
In general reasoning tests, Muse Glimmer excelled in four out of six assessments. On IFBench, it scored 77.0, surpassing Gemma4-31B at 76.0 and Qwen3.6-27B at 70.8. On AIME 2026, Muse Glimmer scored 94.7, compared to 89.2 for Gemma4-31B and 94.1 for Qwen3.6-27B. The model also led on AA-LCR with 80.0, ahead of 68.3 and 73.3.
However, Gemma4-31B took the lead on GPQA Diamond with 85.7, followed by Muse Glimmer at 83.5 and Qwen3.6-27B at 84.2. On Humanity’s Last Exam, Text No Tools, Gemma4-31B scored 23.6, while Muse Glimmer obtained 22.0.
Memory Challenges and Deployment Solutions
Meta designed Muse Glimmer to operate with reduced memory. A full-precision model of 30 billion parameters would require over 55 GB of memory, but through weight quantization to about 4 bits, the model is reduced to less than 20 GB. This allows for memory to be freed up for a KV cache, a perception encoder, and a speculative decoding writer. Meta recommends a memory envelope of 24 GB or 32 GB for these components.
The DFlash-based writer proposes verified token blocks processed in parallel by the main model, thereby accelerating generation while maintaining high output quality. Although Meta has not provided specific figures on performance in terms of tokens per second or energy consumption, tests on hardware like MacBook M4-Max, MacBook M5-Max, and RTX-5090 have shown smooth, real-time interaction with the agent.
The model weights are available on Hugging Face, and Meta plans integrations with llama.cpp, MLX, and ExecuTorch in the near future.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.