GPUFits

Models / Gemma 3 27B

Gemma 3 27B VRAM Requirements & GPU Pairing

Gemma 3 27B is Google's dense flagship distilled from Gemini: 27.4B parameters, 62 layers, a native 128K context, and vision input, with English quality and instruction following in the first tier of the 30B class. The hardware catch is its KV cache: 62 layers × 16 KV heads make the 8K fp16 KV a hefty ~4.2GB — two to three times its peers. Q4_K_M @8K totals ~22.4GB, which is a tight fit, not a comfortable one, on 24GB cards.

Long context blows VRAM quickly: at 32K the KV alone is ~16.6GB, so beyond 16K context strongly consider q8 KV quantization (halving it). The custom Gemma license requires reading usage restrictions before commercial use — less worry-free than Apache-2.0. Against Mistral Small 24B: slightly better quality, tighter VRAM, stricter license. On a 24GB card wanting long-context headroom, Mistral is safer; Gemma's real comfort zone is the 32GB RTX 5090.

Architecture Specs

Total parameters27.4B
Active parameters 27.4B
Layers62
KV heads 16
Head dim128
Native context128K
Licensegemma
KV bytes per token (fp16)496.0 KB

Gemma 3's 62-layer × 16-KV-head design makes its KV cache significantly larger than peers — quantize KV to q8 for long-context work.

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 29.1 GB 34.8 GB
Q6_K 22.5 GB 28.2 GB
Q4_K_M 16.8 GB 22.4 GB
Q3_K_M 13.7 GB 19.4 GB
Q2_K 10.9 GB 16.5 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Not feasible needs 2 cards Try it →
RTX 3090 24.0 GB Tight fit Q4_K_M ≈42 tok/s Try it →
RTX 4070 Ti Super 16.0 GB Not feasible needs 2 cards Try it →
RTX 4090 24.0 GB Tight fit Q4_K_M ≈45 tok/s Try it →
RTX 5090 32.0 GB Comfortable Q6_K ≈60 tok/s Try it →
RTX A6000 48.0 GB Comfortable Q8_0 ≈20 tok/s Try it →
A100 80GB 80.0 GB Comfortable FP16 ≈28 tok/s Try it →
H100 80GB 80.0 GB Comfortable FP16 ≈46 tok/s Try it →
RX 7900 XTX 24.0 GB Tight fit Q4_K_M ≈43 tok/s Try it →
Mac mini M4 Pro (48GB) 36.0 GB Comfortable Q8_0 ≈7 tok/s Try it →
Mac Studio M4 Max (64GB) 48.0 GB Comfortable Q8_0 ≈14 tok/s Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Comfortable FP16 ≈11 tok/s Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

Can a 24GB GPU run Gemma 3 27B?
Yes, but tight: Q4_K_M @8K is ~22.4GB with under 2GB spare. Use q8 KV quantization and keep context under 16K; for long context, step up to a 32GB card (RTX 5090).
Why is Gemma 3 27B's KV cache so large?
62 layers × 16 KV heads: ~508KB per token in fp16 — nearly 2× Qwen3 32B and over 3× Mistral Small 24B. That's 4.2GB at 8K, and a theoretical ~67GB at full 128K.
Is Gemma 3 27B usable commercially?
Conditionally: the Gemma Terms of Use ban listed use cases and require propagating restrictions downstream. Read them before commercial deployment — it is less carefree than Apache-2.0 models.

Related guides

Try Gemma 3 27B in the GPU compatibility checker →

Data verified 2026-09-01