Models / Muse Glimmer 30B
Muse Glimmer 30B VRAM Requirements & GPU Pairing
Muse Glimmer 30B (released August 2026) takes a different hybrid route: in every 4 of its 52 layers, 3 use 2048-token sliding-window attention and 1 is global, with only 2 KV head groups — so the KV cache is tiny by design, ~0.44GB at 8K (our standard formula still overestimates it, same as the gpt-oss sliding-window precedent). The architecture includes a ~1.8B ViT vision encoder for native multimodality; Apache-2.0, 128K context.
Q4_K_M @8K totals ~20.1GB — tight on a single 24GB card, with Q6_K reachable on 32GB. Long context adds almost no VRAM pressure, a structural advantage over Gemma 3 27B. The sliding-window trade-off matches gpt-oss: ultra-long-range pinpoint citation is weaker than full-attention models. The sweet spot is a multimodal long-document assistant on a 24GB card — PDFs, screenshots, long meeting notes. Pure-text quality trades blows with Qwen3 32B, but it is far friendlier on both VRAM and speed.
Architecture Specs
| Total parameters | 29.6B |
| Active parameters | 29.6B |
| Layers | 52 |
| KV heads | 2 |
| Head dim | 128 |
| Native context | 128K |
| License | apache-2.0 |
| KV bytes per token (fp16) | 52.0 KB |
Sliding-window note: 3/4 of layers use 2048-token windows. The table computes KV with the standard full-attention formula; real usage is lower (see the gpt-oss precedent), so verdicts are conservative.
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 31.5 GB | 33.4 GB |
| Q6_K | 24.3 GB | 26.2 GB |
| Q4_K_M | 18.1 GB | 20.1 GB |
| Q3_K_M | 14.8 GB | 16.7 GB |
| Q2_K | 11.7 GB | 13.7 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Not feasible | needs 2 cards | — | Try it → |
| RTX 3090 | 24.0 GB | Tight fit | Q5_K_M | ≈33 tok/s | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Not feasible | Q2_K | ≈43 tok/s | Try it → |
| RTX 4090 | 24.0 GB | Tight fit | Q5_K_M | ≈36 tok/s | Try it → |
| RTX 5090 | 32.0 GB | Comfortable | Q6_K | ≈55 tok/s | Try it → |
| RTX A6000 | 48.0 GB | Comfortable | Q8_0 | ≈18 tok/s | Try it → |
| A100 80GB | 80.0 GB | Comfortable | FP16 | ≈26 tok/s | Try it → |
| H100 80GB | 80.0 GB | Comfortable | FP16 | ≈42 tok/s | Try it → |
| RX 7900 XTX | 24.0 GB | Tight fit | Q5_K_M | ≈34 tok/s | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Comfortable | Q8_0 | ≈7 tok/s | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Comfortable | Q8_0 | ≈13 tok/s | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Comfortable | FP16 | ≈10 tok/s | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- How much VRAM does Muse Glimmer 30B need?
- Q4_K_M @8K totals ~20.1GB — tight on 24GB; Q6_K (~26.2GB) needs 32GB. KV cache is tiny (~0.44GB at 8K), so long-context growth is far below full-attention models.
- Does sliding-window attention hurt quality?
- Barely for most tasks — 3/4 of layers see only the last 2048 tokens while 1/4 see everything, the same battle-tested design as gpt-oss. What suffers is pinpoint citation across extreme distances.
- Muse Glimmer 30B or Qwen3 32B?
- Pick Muse Glimmer for multimodality, low long-context VRAM, and slightly faster decode (2 KV heads + sliding window). Pick Qwen3 32B for pure-text dense quality and the mature Qwen ecosystem, accepting a larger KV and slower decode.
Related guides
- KV Cache Explained: Why Long Contexts Eat Your VRAM
- The VRAM Cost of Long Context (and How to Right-Size It)
- Quantization Explained: Q4 vs Q8 and What You Actually Lose
Try Muse Glimmer 30B in the GPU compatibility checker →
Data verified 2026-09-01