GPUs / RTX A6000
RTX A6000 for Local LLMs: What It Runs and How Fast
The RTX A6000 is the only standard answer to “run 70B on one card”: 48GB of ECC VRAM and 768 GB/s bandwidth — Llama 3.3 70B Q4 (~47.4GB) barely fits, and Qwen3 32B can even attempt Q8_0 (~38.5GB, tight). Used prices sit around $2,500-3,200 (August 2026), more than double the per-GB cost of 24GB cards — you are buying capacity and certainty, not bandwidth.
Blower single-slot cooling, 300W TDP, and NVLink support make it a natural multi-GPU workstation card: two give you 96GB, where gpt-oss-120b Q3 (~60GB) is comfortable. Accept the downsides: Ampere is old, 768 GB/s means a theoretical ~13 tok/s on 70B, and ECC adds minor overhead in some frameworks. It suits professionals who want big-VRAM certainty without multi-GPU fiddling or the Apple ecosystem; the budget alternative at the same capacity is dual 3090s — at the cost of PCIe layer-splitting complexity and double the power draw.
Specs
| Nominal VRAM | 48 GB |
| Usable VRAM (for models) | 48.0 GB |
| Memory bandwidth | 768 GB/s |
| Type | Discrete GPU |
| Release year | 2020 |
| MSRP | $4,650 |
| Used reference price | $2,800 |
| TDP | 300 W |
Value Metrics
Based on used reference price; used prices are market estimates, see commentary for volatility.
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | FP16 | ≈90 tok/s | Try it → |
| Llama 3.1 8B | 7.5 GB | Comfortable | FP16 | ≈36 tok/s | Try it → |
| Qwen3 8B | 7.7 GB | Comfortable | FP16 | ≈35 tok/s | Try it → |
| Phi-4 14B | 12.2 GB | Comfortable | FP16 | ≈20 tok/s | Try it → |
| Mistral Small 3.2 24B | 17.5 GB | Comfortable | Q8_0 | ≈23 tok/s | Try it → |
| Gemma 3 27B | 22.4 GB | Comfortable | Q8_0 | ≈20 tok/s | Try it → |
| Qwen3.8 27B | 20.7 GB | Comfortable | Q8_0 | ≈19 tok/s | Try it → |
| Muse Glimmer 30B | 20.1 GB | Comfortable | Q8_0 | ≈18 tok/s | Try it → |
| Qwen3 30B-A3B | 21.0 GB | Comfortable | Q8_0 | ≈164 tok/s | Try it → |
| Qwen3 32B | 23.7 GB | Comfortable | Q8_0 | ≈17 tok/s | Try it → |
| gpt-oss-20b | 14.7 GB | Comfortable | FP16 | ≈80 tok/s | Try it → |
| Llama 3.3 70B | 47.4 GB | Tight fit | Q4_K_M | ≈13 tok/s | Try it → |
| gpt-oss-120b | 73.6 GB | Not feasible | needs 2 cards | — | Try it → |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- Can the RTX A6000 48GB run 70B?
- Yes, tightly: Llama 3.3 70B Q4_K_M @8K is ~47.4GB — it just fits, at a theoretical ~13 tok/s. For headroom or Q6/Q8, you need an A100 80GB or two cards.
- A6000 or dual 3090s (48GB)?
- Same capacity: dual 3090s offer more aggregate bandwidth (936×2×0.85 ≈ 1591 vs 768 GB/s) for less money, but demand PCIe layer-splitting and 700W. The A6000 is quiet, deterministic, NVLink-capable. Convenience: A6000; value: dual 3090s.
- What is a fair used price for an A6000?
- Around $2,500-3,200 (est., Aug 2026). Above $3,200, look at a Mac Studio M4 Max 64GB ($2,899 new with warranty) or the dual-3090 route instead.
Related guides
- Multi-GPU Setup for Local LLMs: NVLink, PCIe & Layer Splitting
- Best GPU for Local LLMs in 2026: Every Budget Tier
- Why Memory Bandwidth Determines LLM Inference Speed
Try RTX A6000 in the GPU compatibility checker →
Data verified 2026-09-01