GPUFits

GPU Value Rankings: Which Card Is Actually Worth Your Money for Local LLMs?

Most "best GPU" lists just repeat spec sheets. These rankings are different: every number below is computed at build time from our own database of 12 consumer, workstation, and Apple unified-memory GPUs. The methodology:

  • Price: used-market price first; cards without used data (RTX 5090, A100, H100, Apple machines) use MSRP or a market reference. Used prices follow eBay sold/listed listings and are updated monthly.
  • Speed: a theoretical estimate — bandwidth × 0.75 ÷ bytes of weights read per token. Real-world results vary with framework, drivers, and CPU; expect roughly ±30%.
  • VRAM: Apple unified memory is counted at 75% usable (it is shared with the OS and apps); discrete GPUs use their rated capacity.
  • Data verified: 2026-09-01.

Board 1 · Memory Bandwidth per Dollar

Memory bandwidth sets the speed ceiling for LLM inference. Formula: bandwidth (GB/s) ÷ price.

# GPU Bandwidth Price GB/s per $
1 RX 7900 XTX 960 GB/s $650 1.48
2 RTX 3060 12GB 360 GB/s $260vs last month +44% 1.38
3 RTX 5090 1,792 GB/s $1,999 0.90
4 RTX 4070 Ti Super 672 GB/s $760vs last month +38% 0.88
5 RTX 3090 936 GB/s $1,100vs last month +10% 0.85
6 RTX 4090 1,008 GB/s $2,400vs last month +118% 0.42
7 RTX A6000 768 GB/s $2,800 0.27
8 Mac Studio M4 Max (64GB) 546 GB/s $2,899 0.19
9 Mac Studio M3 Ultra (96GB) 819 GB/s $5,299 0.15
10 Mac mini M4 Pro (48GB) 273 GB/s $1,799 0.15
11 A100 80GB 2,039 GB/s $15,000 0.14
12 H100 80GB 3,350 GB/s $30,000 0.11

The RX 7900 XTX takes first place at roughly $650 used — 1.48 GB/s per dollar makes it the bandwidth sweet spot. The RTX 3060 12GB holds second at 1.39 even after the DRAM shortage pushed its used price up ~40% (from $170-220 in July to $250-275, with NVIDIA relaunching it at retail for $339). Even a price hike couldn't dethrone it.

The RTX 5090 ranks third at its $1,999 MSRP — but with secondary-market prices at about 2.1× MSRP ($4,288-4,381 new), it drops out of the top spots at real street prices. It's only a deal if you can actually buy it at MSRP. The A100 and H100 at the bottom prove one thing: datacenter bandwidth costs 10× more per unit. Individuals running inference shouldn't look up to them.

Board 2 · VRAM per Dollar

VRAM decides how large a model you can load. Formula: usable VRAM ÷ price (Apple unified memory at 75%).

# GPU Usable VRAM Price GB per $100
1 RTX 3060 12GB 12.0 GB $260vs last month +44% 4.62
2 RX 7900 XTX 24.0 GB $650 3.69
3 RTX 3090 24.0 GB $1,100vs last month +10% 2.18
4 RTX 4070 Ti Super 16.0 GB $760vs last month +38% 2.11
5 Mac mini M4 Pro (48GB)(×0.75) 36.0 GB $1,799 2.00
6 RTX A6000 48.0 GB $2,800 1.71
7 Mac Studio M4 Max (64GB)(×0.75) 48.0 GB $2,899 1.66
8 RTX 5090 32.0 GB $1,999 1.60
9 Mac Studio M3 Ultra (96GB)(×0.75) 72.0 GB $5,299 1.36
10 RTX 4090 24.0 GB $2,400vs last month +118% 1.00
11 A100 80GB 80.0 GB $15,000 0.53
12 H100 80GB 80.0 GB $30,000 0.27

The RTX 3060 12GB leads by a wide margin at 4.6GB per $100 — which is exactly why it has topped entry-level recommendations for years. The RX 7900 XTX is second at 3.7GB; a 24GB card that launched at $999 and now sells for $650 used remains AMD's classic capacity-per-dollar play.

The RTX 3090 takes third at 2.2GB/$100 despite a persistent 24GB premium ($1,095-1,181 used), and it is the workhorse of budget multi-GPU 70B builds. The adjusted Mac mini M4 Pro 48GB (36GB usable) breaks into the top five — if you want a quiet, low-power box that swallows 30B-class models, it is the only answer at $1,799.

Board 3 · 8B Model Speed per Dollar

Benchmarked against Llama 3.1 8B Q4_K_M (4.9GB of weights read per token). Formula: theoretical tok/s ÷ price. Estimate, ±30%.

# GPU Est. speed Price tok/s per $100
1 RX 7900 XTX ~146 tok/s $650 22.5
2 RTX 3060 12GB ~55 tok/s $260vs last month +44% 21.1
3 RTX 5090 ~273 tok/s $1,999 13.7
4 RTX 4070 Ti Super ~102 tok/s $760vs last month +38% 13.5
5 RTX 3090 ~143 tok/s $1,100vs last month +10% 13.0
6 RTX 4090 ~154 tok/s $2,400vs last month +118% 6.4
7 RTX A6000 ~117 tok/s $2,800 4.2
8 Mac Studio M4 Max (64GB) ~83 tok/s $2,899 2.9
9 Mac Studio M3 Ultra (96GB) ~125 tok/s $5,299 2.4
10 Mac mini M4 Pro (48GB) ~42 tok/s $1,799 2.3
11 A100 80GB ~311 tok/s $15,000 2.1
12 H100 80GB ~511 tok/s $30,000 1.7

For 8B-class models, the RX 7900 XTX and RTX 3060 12GB are nearly tied on speed per dollar (22.5 vs 21.1 tok/s per $100) — the sweet spots of the $650 and $260 budgets respectively. The former is almost three times as fast; the latter costs a fraction. The RTX 5090 is third at 13.7: its absolute 273 tok/s is the fastest of any consumer card here, but the price dilutes the win.

The cautionary tale is the RTX 4090: discontinued, with used prices anchored to the 5090 and climbing to $2,350-2,500. Its speed per dollar is a third of the 3060's — buying one just for 8B models means every dollar goes to the discontinuation premium.

Board 4 · Cheapest Ways to Run a 70B Model

The cheapest setups that comfortably fit Llama 3.3 70B Q4_K_M at 8K context (47.4GB total = 43.2 weights + 2.7 KV cache + 1.5 overhead), sorted by total price. Cards that can't fit it alone show the required card count; total = unit price × cards (motherboard/PSU/cooling not included). Unified-memory Macs are not stacked.

# Setup Cards Total price Headroom
1 RTX 3060 12GBneeds 4 cards 4 $1,040vs last month +44% Tight fit
2 RX 7900 XTXneeds 2 cards 2 $1,300 Tight fit
3 RTX 3090needs 2 cards 2 $2,200vs last month +10% Tight fit
4 RTX 4070 Ti Superneeds 3 cards 3 $2,280vs last month +38% Tight fit
5 RTX A6000 1 $2,800 Tight fit
6 Mac Studio M4 Max (64GB) 1 $2,899 Tight fit
7 RTX 5090needs 2 cards 2 $3,998 Comfortable
8 RTX 4090needs 2 cards 2 $4,800vs last month +118% Tight fit
9 Mac Studio M3 Ultra (96GB) 1 $5,299 Comfortable
10 A100 80GB 1 $15,000 Comfortable
11 H100 80GB 1 $30,000 Comfortable

The cheapest ticket into 70B territory is 4× RTX 3060 12GB ($1,040) — but it's a tight fit, and four cards bring hidden costs in power delivery, cooling, and PCIe slots. The real sweet spot is 2× RX 7900 XTX ($1,300): 48GB across two cards for only $260 more, with a fraction of the engineering pain. A step up, 2× RTX 3090 ($2,200) buys more bandwidth and a smoother CUDA ecosystem.

Prefer a single card? The RTX A6000 48GB ($2,800) is the cheapest single-card answer (tight); 2× RTX 5090 ($3,998) is the lowest-priced "comfortable" setup; and the Mac Studio M3 Ultra 96GB ($5,299) is the only single desktop machine that runs 70B comfortably — at the cost of a ~125 tok/s theoretical ceiling. Silent and power-sipping, but not fast.

Keep researching

FAQ

Are these rankings based on new or used prices?
Used prices take priority: cards that have been on the market for two years or more are ranked by their eBay sold/listed used price. Cards with no used-price data (RTX 5090, A100, H100, and the Apple machines) fall back to MSRP or a market reference price. Used prices are updated monthly; current data was verified on 2026-09-01.
Does the best bandwidth per dollar also mean the fastest model inference?
Not necessarily. Your VRAM has to fit the model weights plus KV cache first — if it doesn't, you end up lowering the quantization, stacking multiple cards, or offloading to the CPU, which destroys speed. Bandwidth only sets the tok/s ceiling; VRAM decides whether you can reach it. Check the VRAM and 70B-threshold boards first, then look at bandwidth and speed.
Why is Apple unified memory counted at 75%?
Apple Silicon shares its unified memory with macOS and every running app, so the GPU-usable ceiling is roughly 75% of total memory (the recommended working set). A Mac mini M4 Pro 48GB therefore competes at 36GB, and a Mac Studio M3 Ultra 96GB at 72GB. Discrete GPUs use their full rated VRAM.

Prices and specs verified 2026-09-01; used prices updated monthly. Every ranking is computed from data at build time, never hand-entered.