GPUFits

GPU Comparisons / RTX 4090 vs RTX 5090

RTX 4090 vs RTX 5090 for Local LLMs: Which Should You Pick

Verdict

If you can buy at MSRP, take the 5090: the 32B tier moves from tight to comfortable and decode is 78% faster. At street prices neither card is rational — the 4090 carries a discontinuation premium, the 5090 trades at ~2× MSRP.

Spec Comparison

RTX 4090 RTX 5090
Nominal VRAM24 GB32 GB
Usable VRAM (for models)24.0 GB32.0 GB
Memory bandwidth1008 GB/s1792 GB/s
TypeDiscrete GPUDiscrete GPU
Release year20222025
MSRP$1,599$1,999
Used reference price $2,400
TDP450 W575 W

Model Verdict Matrix (Q4_K_M @8K)

Model VRAM needed RTX 4090 Theoretical speed RTX 5090 Theoretical speed
Llama 3.2 3B 4.4 GB Comfortable ≈385 Comfortable ≈684
Llama 3.1 8B 7.5 GB Comfortable ≈154 Comfortable ≈273
Qwen3 8B 7.7 GB Comfortable ≈151 Comfortable ≈268
Phi-4 14B 12.2 GB Comfortable ≈84 Comfortable ≈149
Mistral Small 3.2 24B 17.5 GB Comfortable ≈51 Comfortable ≈91
Gemma 3 27B 22.4 GB Tight fit ≈45 Comfortable ≈80
Qwen3.8 27B 20.7 GB Tight fit ≈44 Comfortable ≈79
Muse Glimmer 30B 20.1 GB Tight fit ≈42 Comfortable ≈74
Qwen3 30B-A3B 21.0 GB Tight fit ≈374 Comfortable ≈665
Qwen3 32B 23.7 GB Tight fit ≈38 Comfortable ≈67
gpt-oss-20b 14.7 GB Comfortable ≈343 Comfortable ≈610
Llama 3.3 70B 47.4 GB Not feasible Not feasible
gpt-oss-120b 73.6 GB Not feasible Not feasible
DeepSeek-R1 671B 413.1 GB Not feasible Not feasible

Highlighted rows = watershed models where the two cards' verdicts differ. Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%. Speed is hidden for multi-GPU/infeasible verdicts.

Value Comparison

RTX 4090 RTX 5090
Bandwidth per dollar 0.42 GB/s/$ 0.90 GB/s/$
Usable VRAM per dollar 0.010 GB/$ 0.016 GB/$

Pricing basis: RTX 4090 = used reference price; RTX 5090 = MSRP (no used price yet); used prices are market estimates, see analysis for volatility.

Analysis

The 5090's upgrade over the 4090 is twofold. Bandwidth jumps 1008→1792 GB/s (+78%), lifting decode proportionally across the board: 8B Q4 theoretically ~154→~273 tok/s, Qwen3 32B Q4 ~38→~67 tok/s — the 32B tier finally moves from 'readable' to 'fluid chat'. VRAM grows 24→32GB, moving the entire 27-32B class (Gemma 3 27B, Qwen3.8 27B, Muse Glimmer 30B, Qwen3 30B-A3B, Qwen3 32B) from tight to comfortable; Qwen3 32B can even attempt Q6_K (~30.6GB, tight), and Gemma 3 27B's heavy KV cache stops being a long-context worry.

But both prices are distorted: the 4090 is discontinued, with August 2026 used prices of $2,350-2,500 — over 50% above MSRP; the 5090's $1,999 MSRP is a fiction, with new cards actually at $4,288-4,381 and used at $3,590-4,304 (~2.1× MSRP). Hence an ironic outcome: at real transaction prices the 4090's 0.42 GB/s/$ is the worst of the pair, while the 5090's 0.90 looks great — if you can actually buy at MSRP.

Verdict in three cases: at or near MSRP, the 5090 is the only answer for single-card 32B and beats the 4090; at ~$4,300 street price for pure LLMs, two used 3090s ($2,200, 48GB) offer more VRAM at half the cost; existing 4090 owners should not upgrade — the VRAM tier didn't change, only the speed did. Wait for the next capacity jump.

FAQ

How much faster is the RTX 5090 than the 4090 for LLMs?
1792 vs 1008 GB/s — theoretically +78% decode: 8B Q4 ~154→~273 tok/s, Qwen3 32B Q4 ~38→~67 tok/s. The 32GB VRAM also moves the whole 27-32B tier from tight to comfortable.
Can the RTX 5090 run 70B?
Not on one card: 70B Q4 needs ~47.4GB, over 32GB — you need two. The extra 8GB over the 4090 buys quant headroom in the 27-32B tier (Qwen3 32B can attempt Q6_K at ~30.6GB, tight), not a new model class.
Buy a 4090 or a 5090 right now?
Neither is rational at street prices: used 4090s run $2,350-2,500 (discontinuation premium), 5090s actually cost $3,590-4,300 (~2× MSRP). At MSRP, take the 5090; otherwise two used 3090s ($2,200, 48GB) beat either card for pure LLMs.

Related guides

Data verified 2026-09-01