GPU Comparisons / RTX 4090 vs RTX 5090
RTX 4090 vs RTX 5090 for Local LLMs: Which Should You Pick
If you can buy at MSRP, take the 5090: the 32B tier moves from tight to comfortable and decode is 78% faster. At street prices neither card is rational — the 4090 carries a discontinuation premium, the 5090 trades at ~2× MSRP.
Spec Comparison
| RTX 4090 | RTX 5090 | |
|---|---|---|
| Nominal VRAM | 24 GB | 32 GB |
| Usable VRAM (for models) | 24.0 GB | 32.0 GB |
| Memory bandwidth | 1008 GB/s | 1792 GB/s |
| Type | Discrete GPU | Discrete GPU |
| Release year | 2022 | 2025 |
| MSRP | $1,599 | $1,999 |
| Used reference price | $2,400 | — |
| TDP | 450 W | 575 W |
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | RTX 4090 | Theoretical speed | RTX 5090 | Theoretical speed |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | ≈385 | Comfortable | ≈684 |
| Llama 3.1 8B | 7.5 GB | Comfortable | ≈154 | Comfortable | ≈273 |
| Qwen3 8B | 7.7 GB | Comfortable | ≈151 | Comfortable | ≈268 |
| Phi-4 14B | 12.2 GB | Comfortable | ≈84 | Comfortable | ≈149 |
| Mistral Small 3.2 24B | 17.5 GB | Comfortable | ≈51 | Comfortable | ≈91 |
| Gemma 3 27B | 22.4 GB | Tight fit | ≈45 | Comfortable | ≈80 |
| Qwen3.8 27B | 20.7 GB | Tight fit | ≈44 | Comfortable | ≈79 |
| Muse Glimmer 30B | 20.1 GB | Tight fit | ≈42 | Comfortable | ≈74 |
| Qwen3 30B-A3B | 21.0 GB | Tight fit | ≈374 | Comfortable | ≈665 |
| Qwen3 32B | 23.7 GB | Tight fit | ≈38 | Comfortable | ≈67 |
| gpt-oss-20b | 14.7 GB | Comfortable | ≈343 | Comfortable | ≈610 |
| Llama 3.3 70B | 47.4 GB | Not feasible | — | Not feasible | — |
| gpt-oss-120b | 73.6 GB | Not feasible | — | Not feasible | — |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | Not feasible | — |
Highlighted rows = watershed models where the two cards' verdicts differ. Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%. Speed is hidden for multi-GPU/infeasible verdicts.
Value Comparison
| RTX 4090 | RTX 5090 | |
|---|---|---|
| Bandwidth per dollar | 0.42 GB/s/$ | 0.90 GB/s/$ |
| Usable VRAM per dollar | 0.010 GB/$ | 0.016 GB/$ |
Pricing basis: RTX 4090 = used reference price; RTX 5090 = MSRP (no used price yet); used prices are market estimates, see analysis for volatility.
Analysis
The 5090's upgrade over the 4090 is twofold. Bandwidth jumps 1008→1792 GB/s (+78%), lifting decode proportionally across the board: 8B Q4 theoretically ~154→~273 tok/s, Qwen3 32B Q4 ~38→~67 tok/s — the 32B tier finally moves from 'readable' to 'fluid chat'. VRAM grows 24→32GB, moving the entire 27-32B class (Gemma 3 27B, Qwen3.8 27B, Muse Glimmer 30B, Qwen3 30B-A3B, Qwen3 32B) from tight to comfortable; Qwen3 32B can even attempt Q6_K (~30.6GB, tight), and Gemma 3 27B's heavy KV cache stops being a long-context worry.
But both prices are distorted: the 4090 is discontinued, with August 2026 used prices of $2,350-2,500 — over 50% above MSRP; the 5090's $1,999 MSRP is a fiction, with new cards actually at $4,288-4,381 and used at $3,590-4,304 (~2.1× MSRP). Hence an ironic outcome: at real transaction prices the 4090's 0.42 GB/s/$ is the worst of the pair, while the 5090's 0.90 looks great — if you can actually buy at MSRP.
Verdict in three cases: at or near MSRP, the 5090 is the only answer for single-card 32B and beats the 4090; at ~$4,300 street price for pure LLMs, two used 3090s ($2,200, 48GB) offer more VRAM at half the cost; existing 4090 owners should not upgrade — the VRAM tier didn't change, only the speed did. Wait for the next capacity jump.
FAQ
- How much faster is the RTX 5090 than the 4090 for LLMs?
- 1792 vs 1008 GB/s — theoretically +78% decode: 8B Q4 ~154→~273 tok/s, Qwen3 32B Q4 ~38→~67 tok/s. The 32GB VRAM also moves the whole 27-32B tier from tight to comfortable.
- Can the RTX 5090 run 70B?
- Not on one card: 70B Q4 needs ~47.4GB, over 32GB — you need two. The extra 8GB over the 4090 buys quant headroom in the 27-32B tier (Qwen3 32B can attempt Q6_K at ~30.6GB, tight), not a new model class.
- Buy a 4090 or a 5090 right now?
- Neither is rational at street prices: used 4090s run $2,350-2,500 (discontinuation premium), 5090s actually cost $3,590-4,300 (~2× MSRP). At MSRP, take the 5090; otherwise two used 3090s ($2,200, 48GB) beat either card for pure LLMs.
Related guides
- Why Memory Bandwidth Determines LLM Inference Speed
- Best GPU for Local LLMs in 2026: Every Budget Tier
- FP8, NVFP4, or GGUF Quants: Datacenter Formats vs llama.cpp Formats
- Read the full RTX 4090 breakdown →
- Read the full RTX 5090 breakdown →
- Try RTX 4090 in the GPU compatibility checker →
- Try RTX 5090 in the GPU compatibility checker →
Data verified 2026-09-01