GPUs / RTX 3090
RTX 3090 for Local LLMs: What It Runs and How Fast
The RTX 3090 is the hard currency of local LLMs: 24GB VRAM and 936 GB/s bandwidth, still the used-market value king five-plus years after its 2020 launch — around $1,095-1,181 used in August 2026, with new-old-stock listed from $1,630, and the 24GB premium holding. It runs the 24B tier comfortably (Mistral Small Q4 ~17.5GB), barely fits Qwen3 32B (~23.7GB), and two of them layer-split a 70B Q4 — exactly why its price stays firm: it has both numbers that matter for local LLMs, capacity and bandwidth.
The downsides are purely physical: 350W TDP and hot GDDR6X — on used units always check memory thermal pads and mining history. Dual-slot blower variants stack well for multi-GPU; triple-fan models suit single-card builds. It even has NVLink, though inference layer-splitting doesn't need it. Against the RTX 4090: bandwidth is only 7% behind at half the used price — for pure LLM inference the 3090 is the rational buy, and the savings cover half of a second card.
Specs
| Nominal VRAM | 24 GB |
| Usable VRAM (for models) | 24.0 GB |
| Memory bandwidth | 936 GB/s |
| Type | Discrete GPU |
| Release year | 2020 |
| MSRP | $1,499 |
| Used reference price | $1,100vs last month +10% |
| TDP | 350 W |
Value Metrics
Based on used reference price; used prices are market estimates, see commentary for volatility.
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | FP16 | ≈109 tok/s | Try it → |
| Llama 3.1 8B | 7.5 GB | Comfortable | FP16 | ≈44 tok/s | Try it → |
| Qwen3 8B | 7.7 GB | Comfortable | FP16 | ≈43 tok/s | Try it → |
| Phi-4 14B | 12.2 GB | Comfortable | Q8_0 | ≈45 tok/s | Try it → |
| Mistral Small 3.2 24B | 17.5 GB | Comfortable | Q6_K | ≈36 tok/s | Try it → |
| Gemma 3 27B | 22.4 GB | Tight fit | Q4_K_M | ≈42 tok/s | Try it → |
| Qwen3.8 27B | 20.7 GB | Tight fit | Q5_K_M | ≈35 tok/s | Try it → |
| Muse Glimmer 30B | 20.1 GB | Tight fit | Q5_K_M | ≈33 tok/s | Try it → |
| Qwen3 30B-A3B | 21.0 GB | Tight fit | Q4_K_M | ≈347 tok/s | Try it → |
| Qwen3 32B | 23.7 GB | Tight fit | Q4_K_M | ≈35 tok/s | Try it → |
| gpt-oss-20b | 14.7 GB | Comfortable | Q6_K | ≈237 tok/s | Try it → |
| Llama 3.3 70B | 47.4 GB | Not feasible | needs 2 cards | — | Try it → |
| gpt-oss-120b | 73.6 GB | Not feasible | needs 4 cards | — | Try it → |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- What can the RTX 3090 run?
- At Q4_K_M @8K: the 24B tier comfortably, Qwen3 32B tight (~23.7GB), 70B with two cards. The MoE Qwen3 30B-A3B theoretically hits ~347 tok/s on it — an excellent experience.
- How to inspect a used RTX 3090?
- Three priorities: GDDR6X memory temps (check junction temp under load; over 100°C needs new pads), mining history (ask the use case, inspect warranty seals), and fan/power-connector condition. Blower models are multi-GPU friendly but loud.
- RTX 3090 or RTX 4090?
- For pure LLM: the 3090 — 7% less bandwidth at half the used price. Consider the 4090 only for single-card gaming, lower power, and newer architecture — but at ~$2,400 used it has far outrun its LLM value.
Related guides
- Best GPU for Local LLMs in 2026: Every Budget Tier
- Multi-GPU Setup for Local LLMs: NVLink, PCIe & Layer Splitting
- Why Memory Bandwidth Determines LLM Inference Speed
Try RTX 3090 in the GPU compatibility checker →
Data verified 2026-09-01