GPUFits

Models / Qwen3.8 27B

Qwen3.8 27B VRAM Requirements & GPU Pairing

Qwen3.8 27B (released August 2026) has the most radical architecture in this list: only 16 of its 64 layers are standard full attention (24 Q heads / 4 KV heads / headDim 256, one every 4 layers), while the other 48 use Gated DeltaNet linear attention — whose KV footprint does not grow with context length. That is how it natively supports 256K context with almost no VRAM growth. Apache-2.0 licensed; the parameter count includes a vision encoder.

Keep two ledgers for hardware: weights at Q4_K_M are ~17GB — a tight fit on 24GB cards (~20.7GB total). Our table shows ~2.1GB KV at 8K via the standard GQA formula, but only the full-attention layers produce conventional KV — the formula overestimates ~4×, so real long-context feasibility is better than our matrix suggests. The cost of linear attention is slightly weaker long-range pinpoint retrieval (needle tests), fully adequate for everyday RAG and very-long-document summarization. If you want 27B-class quality with truly long context, it is currently the only answer.

Architecture Specs

Total parameters27.8B
Active parameters 27.8B
Layers64
KV heads 4
Head dim256
Native context256K
Licenseapache-2.0
KV bytes per token (fp16)256.0 KB

Hybrid-architecture note: only 16 of 64 layers are full attention; the 48 Gated DeltaNet layers have context-length-independent KV. The table above uses the standard GQA formula and overestimates KV ~4× — real long-context feasibility is better.

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 29.6 GB 33.2 GB
Q6_K 22.8 GB 26.5 GB
Q4_K_M 17.0 GB 20.7 GB
Q3_K_M 13.9 GB 17.5 GB
Q2_K 11.0 GB 14.7 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Not feasible needs 2 cards Try it →
RTX 3090 24.0 GB Tight fit Q5_K_M ≈35 tok/s Try it →
RTX 4070 Ti Super 16.0 GB Not feasible Q2_K ≈46 tok/s Try it →
RTX 4090 24.0 GB Tight fit Q5_K_M ≈38 tok/s Try it →
RTX 5090 32.0 GB Comfortable Q6_K ≈59 tok/s Try it →
RTX A6000 48.0 GB Comfortable Q8_0 ≈19 tok/s Try it →
A100 80GB 80.0 GB Comfortable FP16 ≈28 tok/s Try it →
H100 80GB 80.0 GB Comfortable FP16 ≈45 tok/s Try it →
RX 7900 XTX 24.0 GB Tight fit Q5_K_M ≈36 tok/s Try it →
Mac mini M4 Pro (48GB) 36.0 GB Comfortable Q8_0 ≈7 tok/s Try it →
Mac Studio M4 Max (64GB) 48.0 GB Comfortable Q8_0 ≈14 tok/s Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Comfortable FP16 ≈11 tok/s Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

How much VRAM does Qwen3.8 27B need?
Q4_K_M @8K totals ~20.7GB — tight on 24GB. Thanks to linear attention, scaling to the full 256K context adds far less KV than standard architectures, the essential difference from Gemma 3 27B.
What is Gated DeltaNet linear attention?
An attention variant whose KV footprint is independent of context length: Qwen3.8 27B uses it in 3 of every 4 layers, with 1 standard attention layer. Long-context VRAM pressure drops sharply, at the cost of slightly weaker ultra-long-range pinpoint recall.
Qwen3.8 27B or Gemma 3 27B?
Pick Qwen3.8 for 256K context, low KV VRAM, and Apache-2.0. Pick Gemma for full-attention long-range retrieval and Google's English quality — accepting a tight 24GB fit and KV blowup at long context.

Related guides

Try Qwen3.8 27B in the GPU compatibility checker →

Data verified 2026-09-01