GPUFits

GPUs / RTX 5090

RTX 5090 for Local LLMs: What It Runs and How Fast

The RTX 5090 is a bandwidth monster: 32GB of GDDR7 at 1792 GB/s — 1.78× the 4090 — translating to a theoretical ~273 tok/s on 8B and ~67 tok/s on Qwen3 32B Q4. For the first time the 32B tier gets a fluid-chat consumer experience. The 32GB VRAM makes Qwen3 32B Q4 (~23.7GB) genuinely comfortable, with Q6_K (~30.6GB) a tight option, and Gemma 3 27B's heavy KV cache stops being a problem.

The problems are price and power: the $1,999 MSRP is a legend — August 2026 street prices run $4,288-4,381 new and $3,590-4,304 used, roughly 2.1× MSRP. TDP is 575W, demanding a 1000W-class PSU and serious airflow. There is exactly one buying logic: you want the strongest single-card experience and money is no object. Rational alternatives: two used 3090s (~$2,200, 48GB) or a Mac Studio M3 Ultra ($5,299, 72GB usable unified memory) win the VRAM-capacity war from below and above.

Specs

Nominal VRAM32 GB
Usable VRAM (for models)32.0 GB
Memory bandwidth1792 GB/s
TypeDiscrete GPU
Release year2025
MSRP$1,999
TDP575 W

Value Metrics

Bandwidth per dollar
0.90 GB/s/$
Usable VRAM per dollar
0.016 GB/$

Based on MSRP (no used price yet); used prices are market estimates, see commentary for volatility.

Model Verdict Matrix (Q4_K_M @8K)

Model VRAM needed Verdict Recommended quant Theoretical speed
Llama 3.2 3B 4.4 GB Comfortable FP16 ≈209 tok/s Try it →
Llama 3.1 8B 7.5 GB Comfortable FP16 ≈84 tok/s Try it →
Qwen3 8B 7.7 GB Comfortable FP16 ≈82 tok/s Try it →
Phi-4 14B 12.2 GB Comfortable Q8_0 ≈86 tok/s Try it →
Mistral Small 3.2 24B 17.5 GB Comfortable Q8_0 ≈53 tok/s Try it →
Gemma 3 27B 22.4 GB Comfortable Q6_K ≈60 tok/s Try it →
Qwen3.8 27B 20.7 GB Comfortable Q6_K ≈59 tok/s Try it →
Muse Glimmer 30B 20.1 GB Comfortable Q6_K ≈55 tok/s Try it →
Qwen3 30B-A3B 21.0 GB Comfortable Q6_K ≈496 tok/s Try it →
Qwen3 32B 23.7 GB Comfortable Q6_K ≈50 tok/s Try it →
gpt-oss-20b 14.7 GB Comfortable Q8_0 ≈351 tok/s Try it →
Llama 3.3 70B 47.4 GB Not feasible needs 2 cards Try it →
gpt-oss-120b 73.6 GB Not feasible needs 3 cards Try it →
DeepSeek-R1 671B 413.1 GB Not feasible Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

What can the RTX 5090 32GB run?
At Q4_K_M @8K: the 32B tier comfortably (Qwen3 32B ~23.7GB, Mistral Small 24B ~17.5GB); Gemma 3 27B finally keeps long-context headroom. 70B needs two cards.
How much faster is the RTX 5090 than the 4090?
LLM decode follows bandwidth: 1792 vs 1008 GB/s is theoretically +78%. 8B goes ~154 → ~273 tok/s, 32B ~37 → ~67 tok/s — the 32B tier moves from “readable” to “fluid”.
Is the RTX 5090 worth its current street price?
The pure LLM math doesn't close: ~$4,300 buys two 3090s (48GB) with $2,000 left over. It only makes sense for single-card 32B fluidity + flagship gaming + zero hassle — or if you can actually buy at the $1,999 MSRP.

Related guides

Try RTX 5090 in the GPU compatibility checker →

Data verified 2026-09-01