GPUs / A100 80GB
A100 80GB for Local LLMs: What It Runs and How Fast
The A100 80GB is the salvage pick of the datacenter retirement wave: 80GB HBM2e with 2039 GB/s on the SXM version — gpt-oss-120b Q4 (~73.6GB) barely fits on one card, 70B is comfortable, and Q8_0 (~79GB) is tight. Market reference price is ~$15,000 (est.); there is no official MSRP.
This is not consumer hardware: SXM boards have no display outputs and depend on server forced-air cooling; only the PCIe variant (slightly lower bandwidth) fits workstations. 400W TDP, datacenter-oriented drivers and virtualization. The buyer profile is a small team or lab wanting 80GB single-card certainty for 120B-class MoE or high-precision 70B without paying H100 prices. Individuals should almost always rent instead — hourly A100 rates on Vast.ai/RunPod are far below the opportunity cost of ownership unless you run 24/7. Against H100: 64% less bandwidth at about half the price — arguably better single-user inference value.
Specs
| Nominal VRAM | 80 GB |
| Usable VRAM (for models) | 80.0 GB |
| Memory bandwidth | 2039 GB/s |
| Type | Discrete GPU |
| Release year | 2020 |
| MSRP | $15,000 |
| TDP | 400 W |
Value Metrics
Based on MSRP (no used price yet); used prices are market estimates, see commentary for volatility.
Model Verdict Matrix (Q4_K_M @8K)
| Model | VRAM needed | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| Llama 3.2 3B | 4.4 GB | Comfortable | FP16 | ≈238 tok/s | Try it → |
| Llama 3.1 8B | 7.5 GB | Comfortable | FP16 | ≈95 tok/s | Try it → |
| Qwen3 8B | 7.7 GB | Comfortable | FP16 | ≈93 tok/s | Try it → |
| Phi-4 14B | 12.2 GB | Comfortable | FP16 | ≈52 tok/s | Try it → |
| Mistral Small 3.2 24B | 17.5 GB | Comfortable | FP16 | ≈32 tok/s | Try it → |
| Gemma 3 27B | 22.4 GB | Comfortable | FP16 | ≈28 tok/s | Try it → |
| Qwen3.8 27B | 20.7 GB | Comfortable | FP16 | ≈28 tok/s | Try it → |
| Muse Glimmer 30B | 20.1 GB | Comfortable | FP16 | ≈26 tok/s | Try it → |
| Qwen3 30B-A3B | 21.0 GB | Comfortable | FP16 | ≈232 tok/s | Try it → |
| Qwen3 32B | 23.7 GB | Comfortable | FP16 | ≈23 tok/s | Try it → |
| gpt-oss-20b | 14.7 GB | Comfortable | FP16 | ≈212 tok/s | Try it → |
| Llama 3.3 70B | 47.4 GB | Comfortable | Q8_0 | ≈20 tok/s | Try it → |
| gpt-oss-120b | 73.6 GB | Tight fit | Q4_K_M | ≈490 tok/s | Try it → |
| DeepSeek-R1 671B | 413.1 GB | Not feasible | — | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- What can the A100 80GB run?
- At Q4_K_M @8K: gpt-oss-120b (~73.6GB) tight; 70B comfortable, Q8_0 (~79GB) tight; everything else comfortably. Full DeepSeek-R1 still needs 6+ cards.
- Can an A100 go into a normal desktop?
- SXM: no — no display output, relies on server airflow. PCIe: yes, in a workstation with verified airflow and power; bandwidth differs slightly from SXM, as does pricing.
- Buy or rent an A100?
- By utilization: only near-24/7 inference justifies ownership. Intermittent use, batch jobs, and fine-tuning are far cheaper rented (hourly on Vast.ai/RunPod). Individuals should default to renting.
Related guides
- Multi-GPU Setup for Local LLMs: NVLink, PCIe & Layer Splitting
- FP8, NVFP4, or GGUF Quants: Datacenter Formats vs llama.cpp Formats
- CPU Offload: Is Partial GPU Offloading Worth It?
Try A100 80GB in the GPU compatibility checker →
Data verified 2026-09-01