Models / Qwen3.8 27B
Qwen3.8 27B VRAM Requirements & GPU Pairing
Qwen3.8 27B (released August 2026) has the most radical architecture in this list: only 16 of its 64 layers are standard full attention (24 Q heads / 4 KV heads / headDim 256, one every 4 layers), while the other 48 use Gated DeltaNet linear attention — whose KV footprint does not grow with context length. That is how it natively supports 256K context with almost no VRAM growth. Apache-2.0 licensed; the parameter count includes a vision encoder.
Keep two ledgers for hardware: weights at Q4_K_M are ~17GB — a tight fit on 24GB cards (~20.7GB total). Our table shows ~2.1GB KV at 8K via the standard GQA formula, but only the full-attention layers produce conventional KV — the formula overestimates ~4×, so real long-context feasibility is better than our matrix suggests. The cost of linear attention is slightly weaker long-range pinpoint retrieval (needle tests), fully adequate for everyday RAG and very-long-document summarization. If you want 27B-class quality with truly long context, it is currently the only answer.
Architecture Specs
| Total parameters | 27.8B |
| Active parameters | 27.8B |
| Layers | 64 |
| KV heads | 4 |
| Head dim | 256 |
| Native context | 256K |
| License | apache-2.0 |
| KV bytes per token (fp16) | 256.0 KB |
Hybrid-architecture note: only 16 of 64 layers are full attention; the 48 Gated DeltaNet layers have context-length-independent KV. The table above uses the standard GQA formula and overestimates KV ~4× — real long-context feasibility is better.
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 29.6 GB | 33.2 GB |
| Q6_K | 22.8 GB | 26.5 GB |
| Q4_K_M | 17.0 GB | 20.7 GB |
| Q3_K_M | 13.9 GB | 17.5 GB |
| Q2_K | 11.0 GB | 14.7 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Not feasible | needs 2 cards | — | Try it → |
| RTX 3090 | 24.0 GB | Tight fit | Q5_K_M | ≈35 tok/s | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Not feasible | Q2_K | ≈46 tok/s | Try it → |
| RTX 4090 | 24.0 GB | Tight fit | Q5_K_M | ≈38 tok/s | Try it → |
| RTX 5090 | 32.0 GB | Comfortable | Q6_K | ≈59 tok/s | Try it → |
| RTX A6000 | 48.0 GB | Comfortable | Q8_0 | ≈19 tok/s | Try it → |
| A100 80GB | 80.0 GB | Comfortable | FP16 | ≈28 tok/s | Try it → |
| H100 80GB | 80.0 GB | Comfortable | FP16 | ≈45 tok/s | Try it → |
| RX 7900 XTX | 24.0 GB | Tight fit | Q5_K_M | ≈36 tok/s | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Comfortable | Q8_0 | ≈7 tok/s | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Comfortable | Q8_0 | ≈14 tok/s | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Comfortable | FP16 | ≈11 tok/s | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- How much VRAM does Qwen3.8 27B need?
- Q4_K_M @8K totals ~20.7GB — tight on 24GB. Thanks to linear attention, scaling to the full 256K context adds far less KV than standard architectures, the essential difference from Gemma 3 27B.
- What is Gated DeltaNet linear attention?
- An attention variant whose KV footprint is independent of context length: Qwen3.8 27B uses it in 3 of every 4 layers, with 1 standard attention layer. Long-context VRAM pressure drops sharply, at the cost of slightly weaker ultra-long-range pinpoint recall.
- Qwen3.8 27B or Gemma 3 27B?
- Pick Qwen3.8 for 256K context, low KV VRAM, and Apache-2.0. Pick Gemma for full-attention long-range retrieval and Google's English quality — accepting a tight 24GB fit and KV blowup at long context.
Related guides
- The VRAM Cost of Long Context (and How to Right-Size It)
- KV Cache Explained: Why Long Contexts Eat Your VRAM
- Best GPU for Local LLMs in 2026: Every Budget Tier
Try Qwen3.8 27B in the GPU compatibility checker →
Data verified 2026-09-01