GPUs / RTX 5090
RTX 5090 for Local LLMs: What It Runs and How Fast
The RTX 5090 is a bandwidth monster: 32GB of GDDR7 at 1792 GB/s — 1.78× the 4090 — translating to a theoretical ~273 tok/s on 8B and ~67 tok/s on Qwen3 32B Q4. For the first time the 32B tier gets a fluid-chat consumer experience. The 32GB VRAM makes Qwen3 32B Q4 (~23.7GB) genuinely comfortable, with Q6_K (~30.6GB) a tight option, and Gemma 3 27B's heavy KV cache stops being a problem.
The problems are price and power: the $1,999 MSRP is a legend — August 2026 street prices run $4,288-4,381 new and $3,590-4,304 used, roughly 2.1× MSRP. TDP is 575W, demanding a 1000W-class PSU and serious airflow. There is exactly one buying logic: you want the strongest single-card experience and money is no object. Rational alternatives: two used 3090s (~$2,200, 48GB) or a Mac Studio M3 Ultra ($5,299, 72GB usable unified memory) win the VRAM-capacity war from below and above.
Specs
| Nominal VRAM | 32 GB |
| Usable VRAM (for models) | 32.0 GB |
| Memory bandwidth | 1792 GB/s |
| Type | Discrete GPU |
| Release year | 2025 |
| MSRP | $1,999 |
| TDP | 575 W |
Value Metrics
Based on MSRP (no used price yet); used prices are market estimates, see commentary for volatility.
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | FP16 | ≈209 tok/s | Try it → |
| Llama 3.1 8B | 7.5 GB | Comfortable | FP16 | ≈84 tok/s | Try it → |
| Qwen3 8B | 7.7 GB | Comfortable | FP16 | ≈82 tok/s | Try it → |
| Phi-4 14B | 12.2 GB | Comfortable | Q8_0 | ≈86 tok/s | Try it → |
| Mistral Small 3.2 24B | 17.5 GB | Comfortable | Q8_0 | ≈53 tok/s | Try it → |
| Gemma 3 27B | 22.4 GB | Comfortable | Q6_K | ≈60 tok/s | Try it → |
| Qwen3.8 27B | 20.7 GB | Comfortable | Q6_K | ≈59 tok/s | Try it → |
| Muse Glimmer 30B | 20.1 GB | Comfortable | Q6_K | ≈55 tok/s | Try it → |
| Qwen3 30B-A3B | 21.0 GB | Comfortable | Q6_K | ≈496 tok/s | Try it → |
| Qwen3 32B | 23.7 GB | Comfortable | Q6_K | ≈50 tok/s | Try it → |
| gpt-oss-20b | 14.7 GB | Comfortable | Q8_0 | ≈351 tok/s | Try it → |
| Llama 3.3 70B | 47.4 GB | Not feasible | needs 2 cards | — | Try it → |
| gpt-oss-120b | 73.6 GB | Not feasible | needs 3 cards | — | Try it → |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- What can the RTX 5090 32GB run?
- At Q4_K_M @8K: the 32B tier comfortably (Qwen3 32B ~23.7GB, Mistral Small 24B ~17.5GB); Gemma 3 27B finally keeps long-context headroom. 70B needs two cards.
- How much faster is the RTX 5090 than the 4090?
- LLM decode follows bandwidth: 1792 vs 1008 GB/s is theoretically +78%. 8B goes ~154 → ~273 tok/s, 32B ~37 → ~67 tok/s — the 32B tier moves from “readable” to “fluid”.
- Is the RTX 5090 worth its current street price?
- The pure LLM math doesn't close: ~$4,300 buys two 3090s (48GB) with $2,000 left over. It only makes sense for single-card 32B fluidity + flagship gaming + zero hassle — or if you can actually buy at the $1,999 MSRP.
Related guides
- Why Memory Bandwidth Determines LLM Inference Speed
- FP8, NVFP4, or GGUF Quants: Datacenter Formats vs llama.cpp Formats
- Best GPU for Local LLMs in 2026: Every Budget Tier
Try RTX 5090 in the GPU compatibility checker →
Data verified 2026-09-01