GPUFits

Models / Phi-4 14B

Phi-4 14B VRAM Requirements & GPU Pairing

Phi-4 is an outlier Microsoft built on high-quality synthetic data: 14.7B dense parameters, 40 layers, 10 KV heads, and the freest license in this list (MIT). Its math and science reasoning punches into the 30B class. The weakness is equally clear: a native context of just 16K — the shortest of all 14 models here — which rules out long documents and long conversations outright.

The hardware math is simple: 40 layers × 10 KV heads make the 8K KV cache ~1.7GB, and Q4_K_M totals ~12.2GB. A 16GB card (RTX 4070 Ti Super) fits comfortably; a 12GB RTX 3060 does not fit at all; 24GB cards can step up to Q8_0 (~18.8GB). The target user is clear: local math, competition problems, and logical reasoning with short inputs. For general chat or RAG, Mistral Small 24B is the more balanced peer. Start at Q4_K_M — at 14.7B parameters, sub-Q4 quality loss bites harder than on 8B models.

Architecture Specs

Total parameters14.7B
Active parameters 14.7B
Layers40
KV heads 10
Head dim128
Native context16K
Licensemit
KV bytes per token (fp16)200.0 KB

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 15.6 GB 18.8 GB
Q6_K 12.1 GB 15.3 GB
Q4_K_M 9.0 GB 12.2 GB
Q3_K_M 7.3 GB 10.5 GB
Q2_K 5.8 GB 9.0 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Needs multi-GPU Q3_K_M ≈37 tok/s Try it →
RTX 3090 24.0 GB Comfortable Q8_0 ≈45 tok/s Try it →
RTX 4070 Ti Super 16.0 GB Comfortable Q6_K ≈42 tok/s Try it →
RTX 4090 24.0 GB Comfortable Q8_0 ≈48 tok/s Try it →
RTX 5090 32.0 GB Comfortable Q8_0 ≈86 tok/s Try it →
RTX A6000 48.0 GB Comfortable FP16 ≈20 tok/s Try it →
A100 80GB 80.0 GB Comfortable FP16 ≈52 tok/s Try it →
H100 80GB 80.0 GB Comfortable FP16 ≈85 tok/s Try it →
RX 7900 XTX 24.0 GB Comfortable Q8_0 ≈46 tok/s Try it →
Mac mini M4 Pro (48GB) 36.0 GB Comfortable FP16 ≈7 tok/s Try it →
Mac Studio M4 Max (64GB) 48.0 GB Comfortable FP16 ≈14 tok/s Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Comfortable FP16 ≈21 tok/s Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

How much VRAM does Phi-4 need?
Q4_K_M @8K totals ~12.2GB — comfortable on 16GB cards; a 12GB card cannot fit it with any margin. Q8_0 is ~18.8GB and needs 24GB.
Is Phi-4's 16K context a dealbreaker?
Depends on the job: math problems, single-turn reasoning, and short documents are fine. Long-PDF Q&A, long conversation memory, or codebase analysis is not — use Mistral Small 24B (128K) there.
What is Phi-4's edge over same-size models?
Punching-above-weight math and science reasoning, plus the MIT license's commercial freedom. Weaknesses: short context, and weaker general chat and knowledge breadth than 24B-class models.

Related guides

Try Phi-4 14B in the GPU compatibility checker →

Data verified 2026-09-01