GPUs / RTX 4090
RTX 4090 for Local LLMs: What It Runs and How Fast
The RTX 4090 is the consumer single-card speed king — and a classic case of market failure: discontinued, with used prices around $2,350-2,500 in August 2026 (+9.7% over 3 months), more than 50% above its $1,599 launch MSRP, its price now anchored to the RTX 5090 rather than its own history. 1008 GB/s plus 24GB means 8B Q4 at a theoretical ~154 tok/s and Qwen3 32B Q4 at ~38 tok/s — single-card coverage up to the 32B tier; two cards tackle 70B.
Do the math before buying: a used 4090 costs about two used 3090s, which together give 48GB and similar aggregate bandwidth — for pure LLM inference, dual 3090s win. The 4090's profile is single-card quietness, relatively lower power, and gaming on the side. Its 450W TDP and the 12VHPWR connector's burn history are used-inspection essentials (check terminal discoloration and seating). If the budget allows exactly one card, it still won't be wrong — it just stopped being the rational choice long ago and became the convenient one.
Specs
| Nominal VRAM | 24 GB |
| Usable VRAM (for models) | 24.0 GB |
| Memory bandwidth | 1008 GB/s |
| Type | Discrete GPU |
| Release year | 2022 |
| MSRP | $1,599 |
| Used reference price | $2,400vs last month +118% |
| TDP | 450 W |
Value Metrics
Based on used reference price; used prices are market estimates, see commentary for volatility.
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | FP16 | ≈118 tok/s | Try it → |
| Llama 3.1 8B | 7.5 GB | Comfortable | FP16 | ≈47 tok/s | Try it → |
| Qwen3 8B | 7.7 GB | Comfortable | FP16 | ≈46 tok/s | Try it → |
| Phi-4 14B | 12.2 GB | Comfortable | Q8_0 | ≈48 tok/s | Try it → |
| Mistral Small 3.2 24B | 17.5 GB | Comfortable | Q6_K | ≈38 tok/s | Try it → |
| Gemma 3 27B | 22.4 GB | Tight fit | Q4_K_M | ≈45 tok/s | Try it → |
| Qwen3.8 27B | 20.7 GB | Tight fit | Q5_K_M | ≈38 tok/s | Try it → |
| Muse Glimmer 30B | 20.1 GB | Tight fit | Q5_K_M | ≈36 tok/s | Try it → |
| Qwen3 30B-A3B | 21.0 GB | Tight fit | Q4_K_M | ≈374 tok/s | Try it → |
| Qwen3 32B | 23.7 GB | Tight fit | Q4_K_M | ≈38 tok/s | Try it → |
| gpt-oss-20b | 14.7 GB | Comfortable | Q6_K | ≈256 tok/s | Try it → |
| Llama 3.3 70B | 47.4 GB | Not feasible | needs 2 cards | — | Try it → |
| gpt-oss-120b | 73.6 GB | Not feasible | needs 4 cards | — | Try it → |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- What can the RTX 4090 run?
- At Q4_K_M @8K: the 24B tier comfortably, Qwen3 32B tight, 70B with two cards. MoE models shine: Qwen3 30B-A3B theoretically ~374 tok/s, gpt-oss-20b ~343 tok/s.
- Why do used RTX 4090s cost more than MSRP?
- Discontinued + the 24GB premium + pricing anchored to the RTX 5090 (street price ~2× MSRP). August 2026 used prices of $2,350-2,500 are pure market pricing, unrelated to the $1,599 MSRP.
- What to check on a used 4090?
- The 12VHPWR connector first: look for discoloration/melting on terminals and confirm original cabling; then check junction temps and fans under load. The 450W TDP wants an 850W+ PSU and good airflow.
Related guides
- Multi-GPU Setup for Local LLMs: NVLink, PCIe & Layer Splitting
- Why Memory Bandwidth Determines LLM Inference Speed
- Best GPU for Local LLMs in 2026: Every Budget Tier
Try RTX 4090 in the GPU compatibility checker →
Data verified 2026-09-01