Models / Phi-4 14B
Phi-4 14B VRAM Requirements & GPU Pairing
Phi-4 is an outlier Microsoft built on high-quality synthetic data: 14.7B dense parameters, 40 layers, 10 KV heads, and the freest license in this list (MIT). Its math and science reasoning punches into the 30B class. The weakness is equally clear: a native context of just 16K — the shortest of all 14 models here — which rules out long documents and long conversations outright.
The hardware math is simple: 40 layers × 10 KV heads make the 8K KV cache ~1.7GB, and Q4_K_M totals ~12.2GB. A 16GB card (RTX 4070 Ti Super) fits comfortably; a 12GB RTX 3060 does not fit at all; 24GB cards can step up to Q8_0 (~18.8GB). The target user is clear: local math, competition problems, and logical reasoning with short inputs. For general chat or RAG, Mistral Small 24B is the more balanced peer. Start at Q4_K_M — at 14.7B parameters, sub-Q4 quality loss bites harder than on 8B models.
Architecture Specs
| Total parameters | 14.7B |
| Active parameters | 14.7B |
| Layers | 40 |
| KV heads | 10 |
| Head dim | 128 |
| Native context | 16K |
| License | mit |
| KV bytes per token (fp16) | 200.0 KB |
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 15.6 GB | 18.8 GB |
| Q6_K | 12.1 GB | 15.3 GB |
| Q4_K_M | 9.0 GB | 12.2 GB |
| Q3_K_M | 7.3 GB | 10.5 GB |
| Q2_K | 5.8 GB | 9.0 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Needs multi-GPU | Q3_K_M | ≈37 tok/s | Try it → |
| RTX 3090 | 24.0 GB | Comfortable | Q8_0 | ≈45 tok/s | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Comfortable | Q6_K | ≈42 tok/s | Try it → |
| RTX 4090 | 24.0 GB | Comfortable | Q8_0 | ≈48 tok/s | Try it → |
| RTX 5090 | 32.0 GB | Comfortable | Q8_0 | ≈86 tok/s | Try it → |
| RTX A6000 | 48.0 GB | Comfortable | FP16 | ≈20 tok/s | Try it → |
| A100 80GB | 80.0 GB | Comfortable | FP16 | ≈52 tok/s | Try it → |
| H100 80GB | 80.0 GB | Comfortable | FP16 | ≈85 tok/s | Try it → |
| RX 7900 XTX | 24.0 GB | Comfortable | Q8_0 | ≈46 tok/s | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Comfortable | FP16 | ≈7 tok/s | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Comfortable | FP16 | ≈14 tok/s | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Comfortable | FP16 | ≈21 tok/s | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- How much VRAM does Phi-4 need?
- Q4_K_M @8K totals ~12.2GB — comfortable on 16GB cards; a 12GB card cannot fit it with any margin. Q8_0 is ~18.8GB and needs 24GB.
- Is Phi-4's 16K context a dealbreaker?
- Depends on the job: math problems, single-turn reasoning, and short documents are fine. Long-PDF Q&A, long conversation memory, or codebase analysis is not — use Mistral Small 24B (128K) there.
- What is Phi-4's edge over same-size models?
- Punching-above-weight math and science reasoning, plus the MIT license's commercial freedom. Weaknesses: short context, and weaker general chat and knowledge breadth than 24B-class models.
Related guides
- Quantization Explained: Q4 vs Q8 and What You Actually Lose
- The VRAM Cost of Long Context (and How to Right-Size It)
- Run GGUF Models Locally: llama.cpp and Ollama Walkthrough
Try Phi-4 14B in the GPU compatibility checker →
Data verified 2026-09-01