GPUs / RTX 4070 Ti Super
RTX 4070 Ti Super for Local LLMs: What It Runs and How Fast
The RTX 4070 Ti Super is the awkward honor student of the 16GB tier: 672 GB/s bandwidth is 28% below the 3090, but with 285W TDP, a January 2024 release, and used prices around $720-800 (~$700 on eBay trackers), it is one of the few big-VRAM N-cards that is essentially mining-free. Its 16GB sweet spot and ceiling are both clear: 8B models fit even at Q8_0 (~11.4GB), gpt-oss-20b Q4 (~14.7GB) is a tight fit, Phi-4 Q4 (~12.2GB) is comfortable — but the entire 24B class is out.
That is the uncrossable gap to 24GB cards: 16GB vs 24GB is not a 50% difference, it is the qualitative line between “can run 32B” and “cannot”. Target buyers: gaming/creator-first users whose LLM use stays in the 8-20B tier, or anyone with zero mining-risk tolerance. If LLM is the main use, a used 3090's 24GB is worth more at the same price. Note 16GB cards have been rising with the memory price surge — comparison-shop before buying.
Specs
| Nominal VRAM | 16 GB |
| Usable VRAM (for models) | 16.0 GB |
| Memory bandwidth | 672 GB/s |
| Type | Discrete GPU |
| Release year | 2024 |
| MSRP | $799 |
| Used reference price | $760vs last month +38% |
| TDP | 285 W |
Value Metrics
Based on used reference price; used prices are market estimates, see commentary for volatility.
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | FP16 | ≈79 tok/s | Try it → |
| Llama 3.1 8B | 7.5 GB | Comfortable | Q8_0 | ≈59 tok/s | Try it → |
| Qwen3 8B | 7.7 GB | Comfortable | Q8_0 | ≈58 tok/s | Try it → |
| Phi-4 14B | 12.2 GB | Comfortable | Q6_K | ≈42 tok/s | Try it → |
| Mistral Small 3.2 24B | 17.5 GB | Needs multi-GPU | Q3_K_M | ≈42 tok/s | Try it → |
| Gemma 3 27B | 22.4 GB | Not feasible | needs 2 cards | — | Try it → |
| Qwen3.8 27B | 20.7 GB | Not feasible | Q2_K | ≈46 tok/s | Try it → |
| Muse Glimmer 30B | 20.1 GB | Not feasible | Q2_K | ≈43 tok/s | Try it → |
| Qwen3 30B-A3B | 21.0 GB | Not feasible | Q2_K | ≈385 tok/s | Try it → |
| Qwen3 32B | 23.7 GB | Not feasible | needs 2 cards | — | Try it → |
| gpt-oss-20b | 14.7 GB | Tight fit | Q4_K_M | ≈229 tok/s | Try it → |
| Llama 3.3 70B | 47.4 GB | Not feasible | needs 3 cards | — | Try it → |
| gpt-oss-120b | 73.6 GB | Not feasible | — | — | Try it → |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- What can the RTX 4070 Ti Super 16GB run?
- At Q4: the 8B tier comfortably, even Q8_0; Phi-4 14B comfortable, gpt-oss-20b tight; the 24B class (Mistral Small / Gemma 3) does not fit at all — that is the hard 16GB ceiling.
- Is 16GB enough for local LLMs?
- Depends on the target tier: fine for 8-14B daily use with higher quants; but only 24GB reaches the 24-32B class — a capacity cliff, not a performance gap. LLM-first users should go 24GB.
- 4070 Ti Super or a used 3090?
- LLM-first: the 3090 — 24GB decides your model tier. Gaming/creator-first with LLM on the side, or zero mining-risk tolerance: the 4070 Ti Super (2024 card, essentially mining-free, 65W lower TDP).
Related guides
- Best GPU for Local LLMs in 2026: Every Budget Tier
- Quantization Explained: Q4 vs Q8 and What You Actually Lose
Try RTX 4070 Ti Super in the GPU compatibility checker →
Data verified 2026-09-01