Skip to content
Main Site News Console

Meta Returns with Muse Glimmer: Local, Agentic, Multimodal, and Open Source

· Hugging Face Translated
教程模型卡

A veteran in the open-source large language model space has brought good news! Muse Glimmer was released today. It is a new multimodal model from Meta, designed especially for local agentic applications. Distilled from Muse, the model has 30 billion (30B) parameters and is released under the Apache 2.0 License. It is ideal for local deployment to protect privacy, reduce costs, or simply for experimentation and development. It targets privacy-conscious use cases such as programming, document analysis, personal assistants, and setups similar to Claw or Hermes.

To celebrate the release, we worked with Meta to provide day-0 support in transformers, llama.cpp, vLLM, Inference Endpoints, and other libraries. We built several interesting applications and share our findings in this article.

Check out the demos below for inspiration.

You can find Muse Glimmer on the Hugging Face Hub.

Benchmarks

Benchmark Results

Scores are listed according to the published results. Bold indicates the best result among the models being compared; ↓ indicates that lower is better.

CategoryBenchmarkMuse Glimmer-30BGemma4-31B
High Reasoning
Qwen3.6-27B
Thinking Mode
General AgentsMCP Atlas75.554.262.5
General AgentsDeepSearch QA74.661.771.1
General Agentsτ³-Banking23.515.116.7
General AgentsWildClawBench47.637.643.2
General AgentsGDPval-AA9538111141
General AgentsGAIA243.336.440.0
General AgentsSkillsBench (with skills)44.332.446.6
General AgentsOSWorld-Verified65.958.575.6
Agentic CodingSWE-Bench Pro51.236.950.2
Agentic CodingSWE-Bench Verified76.066.677.2
Agentic CodingTerminalBench 2.151.743.460.7
Agentic CodingSciCode43.643.439.8
MultimodalCharxiv Reasoning78.877.778.4
MultimodalScreenSpot Pro75.475.976.1
MultimodalOmniDocBench v1.575.872.577.8
MultimodalMMMU Pro747375
SafetyCI MemoriesViolation rate (↓): 26.4
Coverage: 64.8
Violation rate (↓): 12.1
Coverage: 53.0
Violation rate (↓): 53.4
Coverage: 66.9
SafetySiren AgentDojoAttack success rate (↓): 28.4
Utility: 94.2
Attack success rate (↓): 25.6
Utility: 90.8
Attack success rate (↓): 40.3
Utility: 92.7
General Capabilities and ReasoningIFBench77.076.070.8
General Capabilities and ReasoningAIME 202694.789.294.1
General Capabilities and ReasoningGPQA Diamond83.585.784.2
General Capabilities and ReasoningHumanity’s Last Exam (text-only, no tools)22.023.623.1
General Capabilities and ReasoningAA-LCR80.068.373.3
General Capabilities and ReasoningBeam 128K65.158.263.0

Architecture

Muse Glimmer is a dense 30-billion-parameter model consisting of:

  • A 2-billion-parameter ViT-style encoder for vision (Perception Encoder)
  • A 28-billion-parameter text decoder

In addition to the primary VLM, the model includes a speculative decoding drafter implemented using DFlash. This module is optional and can significantly improve generation speed at the cost of some additional memory. We found that this drafter is particularly well suited to structured content generation tasks such as coding.

Text Decoder

The language model uses the following architectural components:

  • Hybrid attention: Three sliding-window layers (with a window size of 2,048 tokens and rotary position embeddings) alternate with a layer using full attention and NoPE (no positional embeddings). The pattern is therefore (SWA, SWA, SWA, Full), repeated 13 times, for a total of 52 layers. This allows the model to use RoPE to preserve relative order and distance information while retaining information globally through NoPE.
  • Gated Grouped-Query Attention: Each key-value head is shared by 16 query heads, reducing the memory footprint of the KV cache by 16× and making generation faster and more cost-efficient.
  • Q-K normalization with additional query scaling: Before computing attention, Muse Glimmer applies RMS normalization to each query and key head to stabilize the attention logits. The model then multiplies the queries by a scaling factor to set the target logit scale after normalization. Additional query scaling behaves similarly to inverse temperature at the softmax layer.

Perception Encoder

Muse Glimmer uses an image encoder to process both images and video. Unlike the relatively small vision encoders found in other VLMs, this is a large 2-billion-parameter ViT-style model designed around the Perception Encoder architecture. Meta previously introduced Perception Encoder as a backbone for a range of downstream spatial and multimodal tasks.

The encoder divides images into patches shaped 2 frames × 3 channels × 14 × 14 and projects them through a linear layer. Absolute positional embeddings are then interpolated from a learned positional table and added to these embeddings. The embeddings are subsequently passed through the vision tower, which consists of 50 layers and GELU MLPs. As with the language model, its attention pattern consists of three window-attention layers followed by one full-attention layer. Within the attention layers, the model applies 2D RoPE to the queries and keys.

After the Transformer, pixel shuffle concatenates adjacent 2×2 groups of spatial tokens, reducing the number of image tokens by 4× without discarding channel information. The merged features are then projected into the shared embedding space of the text decoder.

Video is processed frame by frame through the same encoder. Each frame is converted into patches with the shape [batch, temporal groups, grid height, grid width, 2 frames, 3 channels, 14, 14]. The processor targets a sampling rate of 2 frames per second and limits video clips to a maximum of 96 frames, sampled uniformly throughout the video. It creates timestamped video placeholders and interleaves text with frames, for example, “Time: 0.0s <|video|> x N”; the final video embeddings replace these placeholders immediately before the final projection layer.

Transformers

Upgrade to the latest version of transformers to use Muse Glimmer.

Terminal window
pip install --upgrade transformers accelerate

Muse Glimmer has day-0 support in transformers for both the primary model and the speculative decoding drafter. You can use the AutoModelForMultimodalLM and AutoProcessor classes to load the model and processor.

from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "meta-models/Muse-Glimmer-30B"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)

The same code runs without modification on NVIDIA (CUDA), AMD (ROCm), and Intel (XPU) GPUs; device_map="auto" loads the model onto the available accelerator.

Text-Only Inference

Once the model is loaded, you can run text-only inference as follows.

from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "meta-models/Muse-Glimmer-30B"
# 加载模型
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)
# 提示词
messages = [
{"role": "user", "content": "Write a short joke about saving RAM."},
]
# 处理输入
inputs = processor.apply_chat_template(
messages,
t
#教程#模型卡#AI智能体#开源#Omniapi.co

Published by the 4All API team

Original link:https://huggingface.co/blog/muse-glimmer

Need a mainstream LLM API? 4All API gives you one key to call OpenAI, Anthropic, Google Gemini, Qwen, DeepSeek, and dozens more — at official-pass-through pricing, with enterprise-grade reliability, integrated in 5 minutes.

Sign up for the 4All API console →