GPUFits

GPUs / RTX 4070 Ti Super

RTX 4070 Ti Super for Local LLMs: What It Runs and How Fast

The RTX 4070 Ti Super is the awkward honor student of the 16GB tier: 672 GB/s bandwidth is 28% below the 3090, but with 285W TDP, a January 2024 release, and used prices around $720-800 (~$700 on eBay trackers), it is one of the few big-VRAM N-cards that is essentially mining-free. Its 16GB sweet spot and ceiling are both clear: 8B models fit even at Q8_0 (~11.4GB), gpt-oss-20b Q4 (~14.7GB) is a tight fit, Phi-4 Q4 (~12.2GB) is comfortable — but the entire 24B class is out.

That is the uncrossable gap to 24GB cards: 16GB vs 24GB is not a 50% difference, it is the qualitative line between “can run 32B” and “cannot”. Target buyers: gaming/creator-first users whose LLM use stays in the 8-20B tier, or anyone with zero mining-risk tolerance. If LLM is the main use, a used 3090's 24GB is worth more at the same price. Note 16GB cards have been rising with the memory price surge — comparison-shop before buying.

Specs

Nominal VRAM16 GB
Usable VRAM (for models)16.0 GB
Memory bandwidth672 GB/s
TypeDiscrete GPU
Release year2024
MSRP$799
Used reference price$760vs last month +38%
TDP285 W

Value Metrics

Bandwidth per dollar
0.88 GB/s/$
Usable VRAM per dollar
0.021 GB/$

Based on used reference price; used prices are market estimates, see commentary for volatility.

Model Verdict Matrix (Q4_K_M @8K)

Model VRAM needed Verdict Recommended quant Theoretical speed
Llama 3.2 3B 4.4 GB Comfortable FP16 ≈79 tok/s Try it →
Llama 3.1 8B 7.5 GB Comfortable Q8_0 ≈59 tok/s Try it →
Qwen3 8B 7.7 GB Comfortable Q8_0 ≈58 tok/s Try it →
Phi-4 14B 12.2 GB Comfortable Q6_K ≈42 tok/s Try it →
Mistral Small 3.2 24B 17.5 GB Needs multi-GPU Q3_K_M ≈42 tok/s Try it →
Gemma 3 27B 22.4 GB Not feasible needs 2 cards Try it →
Qwen3.8 27B 20.7 GB Not feasible Q2_K ≈46 tok/s Try it →
Muse Glimmer 30B 20.1 GB Not feasible Q2_K ≈43 tok/s Try it →
Qwen3 30B-A3B 21.0 GB Not feasible Q2_K ≈385 tok/s Try it →
Qwen3 32B 23.7 GB Not feasible needs 2 cards Try it →
gpt-oss-20b 14.7 GB Tight fit Q4_K_M ≈229 tok/s Try it →
Llama 3.3 70B 47.4 GB Not feasible needs 3 cards Try it →
gpt-oss-120b 73.6 GB Not feasible Try it →
DeepSeek-R1 671B 413.1 GB Not feasible Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

What can the RTX 4070 Ti Super 16GB run?
At Q4: the 8B tier comfortably, even Q8_0; Phi-4 14B comfortable, gpt-oss-20b tight; the 24B class (Mistral Small / Gemma 3) does not fit at all — that is the hard 16GB ceiling.
Is 16GB enough for local LLMs?
Depends on the target tier: fine for 8-14B daily use with higher quants; but only 24GB reaches the 24-32B class — a capacity cliff, not a performance gap. LLM-first users should go 24GB.
4070 Ti Super or a used 3090?
LLM-first: the 3090 — 24GB decides your model tier. Gaming/creator-first with LLM on the side, or zero mining-risk tolerance: the 4070 Ti Super (2024 card, essentially mining-free, 65W lower TDP).

Related guides

Try RTX 4070 Ti Super in the GPU compatibility checker →

Data verified 2026-09-01