GPU Value Rankings: Which Card Is Actually Worth Your Money for Local LLMs?
Most "best GPU" lists just repeat spec sheets. These rankings are different: every number below is computed at build time from our own database of 12 consumer, workstation, and Apple unified-memory GPUs. The methodology:
- Price: used-market price first; cards without used data (RTX 5090, A100, H100, Apple machines) use MSRP or a market reference. Used prices follow eBay sold/listed listings and are updated monthly.
- Speed: a theoretical estimate — bandwidth × 0.75 ÷ bytes of weights read per token. Real-world results vary with framework, drivers, and CPU; expect roughly ±30%.
- VRAM: Apple unified memory is counted at 75% usable (it is shared with the OS and apps); discrete GPUs use their rated capacity.
- Data verified: 2026-09-01.
Board 1 · Memory Bandwidth per Dollar
Memory bandwidth sets the speed ceiling for LLM inference. Formula: bandwidth (GB/s) ÷ price.
| # | GPU | Bandwidth | Price | GB/s per $ |
|---|---|---|---|---|
| 1 | RX 7900 XTX | 960 GB/s | $650 | 1.48 |
| 2 | RTX 3060 12GB | 360 GB/s | $260vs last month +44% | 1.38 |
| 3 | RTX 5090 | 1,792 GB/s | $1,999 | 0.90 |
| 4 | RTX 4070 Ti Super | 672 GB/s | $760vs last month +38% | 0.88 |
| 5 | RTX 3090 | 936 GB/s | $1,100vs last month +10% | 0.85 |
| 6 | RTX 4090 | 1,008 GB/s | $2,400vs last month +118% | 0.42 |
| 7 | RTX A6000 | 768 GB/s | $2,800 | 0.27 |
| 8 | Mac Studio M4 Max (64GB) | 546 GB/s | $2,899 | 0.19 |
| 9 | Mac Studio M3 Ultra (96GB) | 819 GB/s | $5,299 | 0.15 |
| 10 | Mac mini M4 Pro (48GB) | 273 GB/s | $1,799 | 0.15 |
| 11 | A100 80GB | 2,039 GB/s | $15,000 | 0.14 |
| 12 | H100 80GB | 3,350 GB/s | $30,000 | 0.11 |
The RX 7900 XTX takes first place at roughly $650 used — 1.48 GB/s per dollar makes it the bandwidth sweet spot. The RTX 3060 12GB holds second at 1.39 even after the DRAM shortage pushed its used price up ~40% (from $170-220 in July to $250-275, with NVIDIA relaunching it at retail for $339). Even a price hike couldn't dethrone it.
The RTX 5090 ranks third at its $1,999 MSRP — but with secondary-market prices at about 2.1× MSRP ($4,288-4,381 new), it drops out of the top spots at real street prices. It's only a deal if you can actually buy it at MSRP. The A100 and H100 at the bottom prove one thing: datacenter bandwidth costs 10× more per unit. Individuals running inference shouldn't look up to them.
Board 2 · VRAM per Dollar
VRAM decides how large a model you can load. Formula: usable VRAM ÷ price (Apple unified memory at 75%).
| # | GPU | Usable VRAM | Price | GB per $100 |
|---|---|---|---|---|
| 1 | RTX 3060 12GB | 12.0 GB | $260vs last month +44% | 4.62 |
| 2 | RX 7900 XTX | 24.0 GB | $650 | 3.69 |
| 3 | RTX 3090 | 24.0 GB | $1,100vs last month +10% | 2.18 |
| 4 | RTX 4070 Ti Super | 16.0 GB | $760vs last month +38% | 2.11 |
| 5 | Mac mini M4 Pro (48GB)(×0.75) | 36.0 GB | $1,799 | 2.00 |
| 6 | RTX A6000 | 48.0 GB | $2,800 | 1.71 |
| 7 | Mac Studio M4 Max (64GB)(×0.75) | 48.0 GB | $2,899 | 1.66 |
| 8 | RTX 5090 | 32.0 GB | $1,999 | 1.60 |
| 9 | Mac Studio M3 Ultra (96GB)(×0.75) | 72.0 GB | $5,299 | 1.36 |
| 10 | RTX 4090 | 24.0 GB | $2,400vs last month +118% | 1.00 |
| 11 | A100 80GB | 80.0 GB | $15,000 | 0.53 |
| 12 | H100 80GB | 80.0 GB | $30,000 | 0.27 |
The RTX 3060 12GB leads by a wide margin at 4.6GB per $100 — which is exactly why it has topped entry-level recommendations for years. The RX 7900 XTX is second at 3.7GB; a 24GB card that launched at $999 and now sells for $650 used remains AMD's classic capacity-per-dollar play.
The RTX 3090 takes third at 2.2GB/$100 despite a persistent 24GB premium ($1,095-1,181 used), and it is the workhorse of budget multi-GPU 70B builds. The adjusted Mac mini M4 Pro 48GB (36GB usable) breaks into the top five — if you want a quiet, low-power box that swallows 30B-class models, it is the only answer at $1,799.
Board 3 · 8B Model Speed per Dollar
Benchmarked against Llama 3.1 8B Q4_K_M (4.9GB of weights read per token). Formula: theoretical tok/s ÷ price. Estimate, ±30%.
| # | GPU | Est. speed | Price | tok/s per $100 |
|---|---|---|---|---|
| 1 | RX 7900 XTX | ~146 tok/s | $650 | 22.5 |
| 2 | RTX 3060 12GB | ~55 tok/s | $260vs last month +44% | 21.1 |
| 3 | RTX 5090 | ~273 tok/s | $1,999 | 13.7 |
| 4 | RTX 4070 Ti Super | ~102 tok/s | $760vs last month +38% | 13.5 |
| 5 | RTX 3090 | ~143 tok/s | $1,100vs last month +10% | 13.0 |
| 6 | RTX 4090 | ~154 tok/s | $2,400vs last month +118% | 6.4 |
| 7 | RTX A6000 | ~117 tok/s | $2,800 | 4.2 |
| 8 | Mac Studio M4 Max (64GB) | ~83 tok/s | $2,899 | 2.9 |
| 9 | Mac Studio M3 Ultra (96GB) | ~125 tok/s | $5,299 | 2.4 |
| 10 | Mac mini M4 Pro (48GB) | ~42 tok/s | $1,799 | 2.3 |
| 11 | A100 80GB | ~311 tok/s | $15,000 | 2.1 |
| 12 | H100 80GB | ~511 tok/s | $30,000 | 1.7 |
For 8B-class models, the RX 7900 XTX and RTX 3060 12GB are nearly tied on speed per dollar (22.5 vs 21.1 tok/s per $100) — the sweet spots of the $650 and $260 budgets respectively. The former is almost three times as fast; the latter costs a fraction. The RTX 5090 is third at 13.7: its absolute 273 tok/s is the fastest of any consumer card here, but the price dilutes the win.
The cautionary tale is the RTX 4090: discontinued, with used prices anchored to the 5090 and climbing to $2,350-2,500. Its speed per dollar is a third of the 3060's — buying one just for 8B models means every dollar goes to the discontinuation premium.
Board 4 · Cheapest Ways to Run a 70B Model
The cheapest setups that comfortably fit Llama 3.3 70B Q4_K_M at 8K context (47.4GB total = 43.2 weights + 2.7 KV cache + 1.5 overhead), sorted by total price. Cards that can't fit it alone show the required card count; total = unit price × cards (motherboard/PSU/cooling not included). Unified-memory Macs are not stacked.
| # | Setup | Cards | Total price | Headroom |
|---|---|---|---|---|
| 1 | RTX 3060 12GBneeds 4 cards | 4 | $1,040vs last month +44% | Tight fit |
| 2 | RX 7900 XTXneeds 2 cards | 2 | $1,300 | Tight fit |
| 3 | RTX 3090needs 2 cards | 2 | $2,200vs last month +10% | Tight fit |
| 4 | RTX 4070 Ti Superneeds 3 cards | 3 | $2,280vs last month +38% | Tight fit |
| 5 | RTX A6000 | 1 | $2,800 | Tight fit |
| 6 | Mac Studio M4 Max (64GB) | 1 | $2,899 | Tight fit |
| 7 | RTX 5090needs 2 cards | 2 | $3,998 | Comfortable |
| 8 | RTX 4090needs 2 cards | 2 | $4,800vs last month +118% | Tight fit |
| 9 | Mac Studio M3 Ultra (96GB) | 1 | $5,299 | Comfortable |
| 10 | A100 80GB | 1 | $15,000 | Comfortable |
| 11 | H100 80GB | 1 | $30,000 | Comfortable |
The cheapest ticket into 70B territory is 4× RTX 3060 12GB ($1,040) — but it's a tight fit, and four cards bring hidden costs in power delivery, cooling, and PCIe slots. The real sweet spot is 2× RX 7900 XTX ($1,300): 48GB across two cards for only $260 more, with a fraction of the engineering pain. A step up, 2× RTX 3090 ($2,200) buys more bandwidth and a smoother CUDA ecosystem.
Prefer a single card? The RTX A6000 48GB ($2,800) is the cheapest single-card answer (tight); 2× RTX 5090 ($3,998) is the lowest-priced "comfortable" setup; and the Mac Studio M3 Ultra 96GB ($5,299) is the only single desktop machine that runs 70B comfortably — at the cost of a ~125 tok/s theoretical ceiling. Silent and power-sipping, but not fast.
Keep researching
- VRAM Calculator — enter a model and a GPU to get exact memory requirements and the recommended quantization
- Local vs API Cost Calculator — work out how quickly local inference pays for itself
- Best GPU for Local LLMs 2026 — the complete buying guide by budget and model size
FAQ
- Are these rankings based on new or used prices?
- Used prices take priority: cards that have been on the market for two years or more are ranked by their eBay sold/listed used price. Cards with no used-price data (RTX 5090, A100, H100, and the Apple machines) fall back to MSRP or a market reference price. Used prices are updated monthly; current data was verified on 2026-09-01.
- Does the best bandwidth per dollar also mean the fastest model inference?
- Not necessarily. Your VRAM has to fit the model weights plus KV cache first — if it doesn't, you end up lowering the quantization, stacking multiple cards, or offloading to the CPU, which destroys speed. Bandwidth only sets the tok/s ceiling; VRAM decides whether you can reach it. Check the VRAM and 70B-threshold boards first, then look at bandwidth and speed.
- Why is Apple unified memory counted at 75%?
- Apple Silicon shares its unified memory with macOS and every running app, so the GPU-usable ceiling is roughly 75% of total memory (the recommended working set). A Mac mini M4 Pro 48GB therefore competes at 36GB, and a Mac Studio M3 Ultra 96GB at 72GB. Discrete GPUs use their full rated VRAM.
Prices and specs verified 2026-09-01; used prices updated monthly. Every ranking is computed from data at build time, never hand-entered.