GPUFits

Can A100 80GB run Mistral Small 3.2 24B?

✅ Comfortable
Mistral Small 3.2 24B @ Q4_K_M · 8K context
Needed
17.5 GB
Usable
80.0 GB

VRAM breakdown

Weights (Q4_K_M)14.7 GB
KV cache (8K context)1.3 GB
Runtime overhead1.5 GB
Total17.5 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

FP16

Estimated generation speed

~32 tok/s @ FP16

Theoretical estimates based on published architecture data and measured GGUF sizes; real-world speed varies ±30%.

Cheaper GPUs that run it

Smaller models this GPU runs well

Other GPUs that run Mistral Small 3.2 24B

FAQ

How much VRAM does Mistral Small 3.2 24B need?
At Q4_K_M with 8K context: 14.7GB weights + 1.3GB KV cache + 1.5GB runtime overhead = 17.5GB total. A100 80GB offers 80.0GB usable VRAM, so the verdict is: Comfortable.
What is the best quantization for Mistral Small 3.2 24B on A100 80GB?
FP16 — the highest tier that still fits within 80.0GB usable VRAM at 8K context. Lower tiers (Q3/Q2) fit too but cost noticeable quality.
How fast does Mistral Small 3.2 24B run on A100 80GB?
About 32 tokens/s at FP16 (theoretical estimate, ±30% in real-world use).

Check another combination in the GPU Checker →

Data verified 2026-09-01