A veteran in the open-source large language model space has brought good news! Muse Glimmer was released today. It is a new multimodal model from Meta, designed especially for local agentic applications. Distilled from Muse, the model has 30 billion (30B) parameters and is released under the Apache 2.0 License. It is ideal for local deployment to protect privacy, reduce costs, or simply for experimentation and development. It targets privacy-conscious use cases such as programming, document analysis, personal assistants, and setups similar to Claw or Hermes.
To celebrate the release, we worked with Meta to provide day-0 support in transformers, llama.cpp, vLLM, Inference Endpoints, and other libraries. We built several interesting applications and share our findings in this article.
Check out the demos below for inspiration.
You can find Muse Glimmer on the Hugging Face Hub.
Benchmarks
Benchmark Results
Scores are listed according to the published results. Bold indicates the best result among the models being compared; ↓ indicates that lower is better.
| Category | Benchmark | Muse Glimmer-30B | Gemma4-31B High Reasoning | Qwen3.6-27B Thinking Mode |
|---|---|---|---|---|
| General Agents | MCP Atlas | 75.5 | 54.2 | 62.5 |
| General Agents | DeepSearch QA | 74.6 | 61.7 | 71.1 |
| General Agents | τ³-Banking | 23.5 | 15.1 | 16.7 |
| General Agents | WildClawBench | 47.6 | 37.6 | 43.2 |
| General Agents | GDPval-AA | 953 | 811 | 1141 |
| General Agents | GAIA2 | 43.3 | 36.4 | 40.0 |
| General Agents | SkillsBench (with skills) | 44.3 | 32.4 | 46.6 |
| General Agents | OSWorld-Verified | 65.9 | 58.5 | 75.6 |
| Agentic Coding | SWE-Bench Pro | 51.2 | 36.9 | 50.2 |
| Agentic Coding | SWE-Bench Verified | 76.0 | 66.6 | 77.2 |
| Agentic Coding | TerminalBench 2.1 | 51.7 | 43.4 | 60.7 |
| Agentic Coding | SciCode | 43.6 | 43.4 | 39.8 |
| Multimodal | Charxiv Reasoning | 78.8 | 77.7 | 78.4 |
| Multimodal | ScreenSpot Pro | 75.4 | 75.9 | 76.1 |
| Multimodal | OmniDocBench v1.5 | 75.8 | 72.5 | 77.8 |
| Multimodal | MMMU Pro | 74 | 73 | 75 |
| Safety | CI Memories | Violation rate (↓): 26.4 Coverage: 64.8 | Violation rate (↓): 12.1 Coverage: 53.0 | Violation rate (↓): 53.4 Coverage: 66.9 |
| Safety | Siren AgentDojo | Attack success rate (↓): 28.4 Utility: 94.2 | Attack success rate (↓): 25.6 Utility: 90.8 | Attack success rate (↓): 40.3 Utility: 92.7 |
| General Capabilities and Reasoning | IFBench | 77.0 | 76.0 | 70.8 |
| General Capabilities and Reasoning | AIME 2026 | 94.7 | 89.2 | 94.1 |
| General Capabilities and Reasoning | GPQA Diamond | 83.5 | 85.7 | 84.2 |
| General Capabilities and Reasoning | Humanity’s Last Exam (text-only, no tools) | 22.0 | 23.6 | 23.1 |
| General Capabilities and Reasoning | AA-LCR | 80.0 | 68.3 | 73.3 |
| General Capabilities and Reasoning | Beam 128K | 65.1 | 58.2 | 63.0 |
Architecture
Muse Glimmer is a dense 30-billion-parameter model consisting of:
- A 2-billion-parameter ViT-style encoder for vision (Perception Encoder)
- A 28-billion-parameter text decoder
In addition to the primary VLM, the model includes a speculative decoding drafter implemented using DFlash. This module is optional and can significantly improve generation speed at the cost of some additional memory. We found that this drafter is particularly well suited to structured content generation tasks such as coding.
Text Decoder
The language model uses the following architectural components:
- Hybrid attention: Three sliding-window layers (with a window size of 2,048 tokens and rotary position embeddings) alternate with a layer using full attention and NoPE (no positional embeddings). The pattern is therefore (SWA, SWA, SWA, Full), repeated 13 times, for a total of 52 layers. This allows the model to use RoPE to preserve relative order and distance information while retaining information globally through NoPE.
- Gated Grouped-Query Attention: Each key-value head is shared by 16 query heads, reducing the memory footprint of the KV cache by 16× and making generation faster and more cost-efficient.
- Q-K normalization with additional query scaling: Before computing attention, Muse Glimmer applies RMS normalization to each query and key head to stabilize the attention logits. The model then multiplies the queries by a scaling factor to set the target logit scale after normalization. Additional query scaling behaves similarly to inverse temperature at the softmax layer.
Perception Encoder
Muse Glimmer uses an image encoder to process both images and video. Unlike the relatively small vision encoders found in other VLMs, this is a large 2-billion-parameter ViT-style model designed around the Perception Encoder architecture. Meta previously introduced Perception Encoder as a backbone for a range of downstream spatial and multimodal tasks.
The encoder divides images into patches shaped 2 frames × 3 channels × 14 × 14 and projects them through a linear layer. Absolute positional embeddings are then interpolated from a learned positional table and added to these embeddings. The embeddings are subsequently passed through the vision tower, which consists of 50 layers and GELU MLPs. As with the language model, its attention pattern consists of three window-attention layers followed by one full-attention layer. Within the attention layers, the model applies 2D RoPE to the queries and keys.
After the Transformer, pixel shuffle concatenates adjacent 2×2 groups of spatial tokens, reducing the number of image tokens by 4× without discarding channel information. The merged features are then projected into the shared embedding space of the text decoder.
Video is processed frame by frame through the same encoder. Each frame is converted into patches with the shape [batch, temporal groups, grid height, grid width, 2 frames, 3 channels, 14, 14]. The processor targets a sampling rate of 2 frames per second and limits video clips to a maximum of 96 frames, sampled uniformly throughout the video. It creates timestamped video placeholders and interleaves text with frames, for example, “Time: 0.0s <|video|> x N”; the final video embeddings replace these placeholders immediately before the final projection layer.
Transformers
Upgrade to the latest version of transformers to use Muse Glimmer.
pip install --upgrade transformers accelerateMuse Glimmer has day-0 support in transformers for both the primary model and the speculative decoding drafter. You can use the AutoModelForMultimodalLM and AutoProcessor classes to load the model and processor.
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "meta-models/Muse-Glimmer-30B"
processor = AutoProcessor.from_pretrained(MODEL_ID)model = AutoModelForMultimodalLM.from_pretrained( MODEL_ID, dtype="auto", device_map="auto")The same code runs without modification on NVIDIA (CUDA), AMD (ROCm), and Intel (XPU) GPUs; device_map="auto" loads the model onto the available accelerator.
Text-Only Inference
Once the model is loaded, you can run text-only inference as follows.
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "meta-models/Muse-Glimmer-30B"
# 加载模型processor = AutoProcessor.from_pretrained(MODEL_ID)model = AutoModelForMultimodalLM.from_pretrained( MODEL_ID, dtype="auto", device_map="auto")
# 提示词messages = [ {"role": "user", "content": "Write a short joke about saving RAM."},]
# 处理输入inputs = processor.apply_chat_template( messages, t