GPUFits

GPU Comparisons / RTX A6000 vs RTX 3090

RTX A6000 vs RTX 3090 for Local LLMs: Which Should You Pick

Verdict

For 70B on one card: RTX A6000 (48GB, fits at the ceiling). If you accept dual-card tinkering: two 3090s — same capacity, double the aggregate bandwidth, ~$600 cheaper.

Spec Comparison

RTX A6000 RTX 3090
Nominal VRAM48 GB24 GB
Usable VRAM (for models)48.0 GB24.0 GB
Memory bandwidth768 GB/s936 GB/s
TypeDiscrete GPUDiscrete GPU
Release year20202020
MSRP$4,650$1,499
Used reference price $2,800 $1,100
TDP300 W350 W

Model Verdict Matrix (Q4_K_M @8K)

Model VRAM needed RTX A6000 Theoretical speed RTX 3090 Theoretical speed
Llama 3.2 3B 4.4 GB Comfortable ≈293 Comfortable ≈357
Llama 3.1 8B 7.5 GB Comfortable ≈117 Comfortable ≈143
Qwen3 8B 7.7 GB Comfortable ≈115 Comfortable ≈140
Phi-4 14B 12.2 GB Comfortable ≈64 Comfortable ≈78
Mistral Small 3.2 24B 17.5 GB Comfortable ≈39 Comfortable ≈48
Gemma 3 27B 22.4 GB Comfortable ≈34 Tight fit ≈42
Qwen3.8 27B 20.7 GB Comfortable ≈34 Tight fit ≈41
Muse Glimmer 30B 20.1 GB Comfortable ≈32 Tight fit ≈39
Qwen3 30B-A3B 21.0 GB Comfortable ≈285 Tight fit ≈347
Qwen3 32B 23.7 GB Comfortable ≈29 Tight fit ≈35
gpt-oss-20b 14.7 GB Comfortable ≈261 Comfortable ≈318
Llama 3.3 70B 47.4 GB Tight fit ≈13 Not feasible
gpt-oss-120b 73.6 GB Not feasible Not feasible
DeepSeek-R1 671B 413.1 GB Not feasible Not feasible

Highlighted rows = watershed models where the two cards' verdicts differ. Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%. Speed is hidden for multi-GPU/infeasible verdicts.

Value Comparison

RTX A6000 RTX 3090
Bandwidth per dollar 0.27 GB/s/$ 0.85 GB/s/$
Usable VRAM per dollar 0.017 GB/$ 0.022 GB/$

Pricing basis: RTX A6000 = used reference price; RTX 3090 = used reference price; used prices are market estimates, see analysis for volatility.

Analysis

The core number of this comparison is 48GB: the A6000 holds Llama 3.3 70B Q4 (~47.4GB) on a single card — tight, at a theoretical ~13 tok/s — while a lone 24GB 3090 marks 70B infeasible in the matrix. But two layer-split 3090s also give 48GB, tying the verdict, and their aggregate bandwidth of 936×2×0.85 ≈ 1591 GB/s is over double the A6000's 768 — 70B at theoretically ~28 tok/s, double the speed too.

The price math: a used A6000 runs $2,500-3,200; dual 3090s cost ~$2,200 — cheaper and faster. What the A6000's premium buys is certainty: single-card blower cooling, 300W TDP (vs 700W combined), ECC memory, NVLink, and no layer-splitting to configure — quiet and stable in a workstation. Dual 3090s demand PCIe lanes on the motherboard, an 850W+ PSU, case airflow, and multi-GPU framework setup, all at once.

Verdict: commercial settings, tinkering-averse, slot- or noise-constrained — the A6000 is the standard answer to '70B on one card'; budget-sensitive, enjoys tinkering, or cares about the speed ceiling — dual 3090s win on value, and day-to-day the two cards can even serve different 24B models independently. Note 70B is at the ceiling on both routes; for Q6/Q8 or gpt-oss-120b, look at the A100 80GB or Mac Studio M3 Ultra 96GB.

FAQ

Can the RTX A6000 run 70B on one card?
Yes, at the ceiling: Llama 3.3 70B Q4_K_M @8K is ~47.4GB against 48GB — a tight fit at a theoretical ~13 tok/s. For Q6/Q8 or long-context headroom you need an 80GB-class card or two GPUs.
Dual 3090s vs A6000 for 70B — which is faster?
Dual 3090s: aggregate bandwidth 936×2×0.85 ≈ 1591 vs 768 GB/s, theoretically ~28 vs ~13 tok/s on 70B, and ~$600 cheaper. The cost is PCIe layer-splitting setup, 700W, and multi-GPU maintenance.
Does the A6000's ECC memory matter for inference?
Marginally: ECC adds negligible overhead in some frameworks but buys stability under sustained load — a real benefit for commercial/24-7 deployments, imperceptible for hobby use.

Related guides

Data verified 2026-09-01