GPUFits

Models / Muse Glimmer 30B

Muse Glimmer 30B VRAM Requirements & GPU Pairing

Muse Glimmer 30B (released August 2026) takes a different hybrid route: in every 4 of its 52 layers, 3 use 2048-token sliding-window attention and 1 is global, with only 2 KV head groups — so the KV cache is tiny by design, ~0.44GB at 8K (our standard formula still overestimates it, same as the gpt-oss sliding-window precedent). The architecture includes a ~1.8B ViT vision encoder for native multimodality; Apache-2.0, 128K context.

Q4_K_M @8K totals ~20.1GB — tight on a single 24GB card, with Q6_K reachable on 32GB. Long context adds almost no VRAM pressure, a structural advantage over Gemma 3 27B. The sliding-window trade-off matches gpt-oss: ultra-long-range pinpoint citation is weaker than full-attention models. The sweet spot is a multimodal long-document assistant on a 24GB card — PDFs, screenshots, long meeting notes. Pure-text quality trades blows with Qwen3 32B, but it is far friendlier on both VRAM and speed.

Architecture Specs

Total parameters29.6B
Active parameters 29.6B
Layers52
KV heads 2
Head dim128
Native context128K
Licenseapache-2.0
KV bytes per token (fp16)52.0 KB

Sliding-window note: 3/4 of layers use 2048-token windows. The table computes KV with the standard full-attention formula; real usage is lower (see the gpt-oss precedent), so verdicts are conservative.

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 31.5 GB 33.4 GB
Q6_K 24.3 GB 26.2 GB
Q4_K_M 18.1 GB 20.1 GB
Q3_K_M 14.8 GB 16.7 GB
Q2_K 11.7 GB 13.7 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Not feasible needs 2 cards Try it →
RTX 3090 24.0 GB Tight fit Q5_K_M ≈33 tok/s Try it →
RTX 4070 Ti Super 16.0 GB Not feasible Q2_K ≈43 tok/s Try it →
RTX 4090 24.0 GB Tight fit Q5_K_M ≈36 tok/s Try it →
RTX 5090 32.0 GB Comfortable Q6_K ≈55 tok/s Try it →
RTX A6000 48.0 GB Comfortable Q8_0 ≈18 tok/s Try it →
A100 80GB 80.0 GB Comfortable FP16 ≈26 tok/s Try it →
H100 80GB 80.0 GB Comfortable FP16 ≈42 tok/s Try it →
RX 7900 XTX 24.0 GB Tight fit Q5_K_M ≈34 tok/s Try it →
Mac mini M4 Pro (48GB) 36.0 GB Comfortable Q8_0 ≈7 tok/s Try it →
Mac Studio M4 Max (64GB) 48.0 GB Comfortable Q8_0 ≈13 tok/s Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Comfortable FP16 ≈10 tok/s Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

How much VRAM does Muse Glimmer 30B need?
Q4_K_M @8K totals ~20.1GB — tight on 24GB; Q6_K (~26.2GB) needs 32GB. KV cache is tiny (~0.44GB at 8K), so long-context growth is far below full-attention models.
Does sliding-window attention hurt quality?
Barely for most tasks — 3/4 of layers see only the last 2048 tokens while 1/4 see everything, the same battle-tested design as gpt-oss. What suffers is pinpoint citation across extreme distances.
Muse Glimmer 30B or Qwen3 32B?
Pick Muse Glimmer for multimodality, low long-context VRAM, and slightly faster decode (2 KV heads + sliding window). Pick Qwen3 32B for pure-text dense quality and the mature Qwen ecosystem, accepting a larger KV and slower decode.

Related guides

Try Muse Glimmer 30B in the GPU compatibility checker →

Data verified 2026-09-01