Models / Gemma 3 27B
Gemma 3 27B VRAM Requirements & GPU Pairing
Gemma 3 27B is Google's dense flagship distilled from Gemini: 27.4B parameters, 62 layers, a native 128K context, and vision input, with English quality and instruction following in the first tier of the 30B class. The hardware catch is its KV cache: 62 layers × 16 KV heads make the 8K fp16 KV a hefty ~4.2GB — two to three times its peers. Q4_K_M @8K totals ~22.4GB, which is a tight fit, not a comfortable one, on 24GB cards.
Long context blows VRAM quickly: at 32K the KV alone is ~16.6GB, so beyond 16K context strongly consider q8 KV quantization (halving it). The custom Gemma license requires reading usage restrictions before commercial use — less worry-free than Apache-2.0. Against Mistral Small 24B: slightly better quality, tighter VRAM, stricter license. On a 24GB card wanting long-context headroom, Mistral is safer; Gemma's real comfort zone is the 32GB RTX 5090.
Architecture Specs
| Total parameters | 27.4B |
| Active parameters | 27.4B |
| Layers | 62 |
| KV heads | 16 |
| Head dim | 128 |
| Native context | 128K |
| License | gemma |
| KV bytes per token (fp16) | 496.0 KB |
Gemma 3's 62-layer × 16-KV-head design makes its KV cache significantly larger than peers — quantize KV to q8 for long-context work.
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 29.1 GB | 34.8 GB |
| Q6_K | 22.5 GB | 28.2 GB |
| Q4_K_M | 16.8 GB | 22.4 GB |
| Q3_K_M | 13.7 GB | 19.4 GB |
| Q2_K | 10.9 GB | 16.5 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Not feasible | needs 2 cards | — | Try it → |
| RTX 3090 | 24.0 GB | Tight fit | Q4_K_M | ≈42 tok/s | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Not feasible | needs 2 cards | — | Try it → |
| RTX 4090 | 24.0 GB | Tight fit | Q4_K_M | ≈45 tok/s | Try it → |
| RTX 5090 | 32.0 GB | Comfortable | Q6_K | ≈60 tok/s | Try it → |
| RTX A6000 | 48.0 GB | Comfortable | Q8_0 | ≈20 tok/s | Try it → |
| A100 80GB | 80.0 GB | Comfortable | FP16 | ≈28 tok/s | Try it → |
| H100 80GB | 80.0 GB | Comfortable | FP16 | ≈46 tok/s | Try it → |
| RX 7900 XTX | 24.0 GB | Tight fit | Q4_K_M | ≈43 tok/s | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Comfortable | Q8_0 | ≈7 tok/s | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Comfortable | Q8_0 | ≈14 tok/s | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Comfortable | FP16 | ≈11 tok/s | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- Can a 24GB GPU run Gemma 3 27B?
- Yes, but tight: Q4_K_M @8K is ~22.4GB with under 2GB spare. Use q8 KV quantization and keep context under 16K; for long context, step up to a 32GB card (RTX 5090).
- Why is Gemma 3 27B's KV cache so large?
- 62 layers × 16 KV heads: ~508KB per token in fp16 — nearly 2× Qwen3 32B and over 3× Mistral Small 24B. That's 4.2GB at 8K, and a theoretical ~67GB at full 128K.
- Is Gemma 3 27B usable commercially?
- Conditionally: the Gemma Terms of Use ban listed use cases and require propagating restrictions downstream. Read them before commercial deployment — it is less carefree than Apache-2.0 models.
Related guides
- KV Cache Explained: Why Long Contexts Eat Your VRAM
- The VRAM Cost of Long Context (and How to Right-Size It)
- Quantization Explained: Q4 vs Q8 and What You Actually Lose
Try Gemma 3 27B in the GPU compatibility checker →
Data verified 2026-09-01