GPUFits

Can RTX 5090 run Mistral Small 3.2 24B?

✅ Comfortable
Mistral Small 3.2 24B @ Q4_K_M · 8K context
Needed
17.5 GB
Usable
32.0 GB

VRAM breakdown

Weights (Q4_K_M)14.7 GB
KV cache (8K context)1.3 GB
Runtime overhead1.5 GB
Total17.5 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

Q8_0

Estimated generation speed

~53 tok/s @ Q8_0

Theoretical estimates based on published architecture data and measured GGUF sizes; real-world speed varies ±30%.

Cheaper GPUs that run it

Smaller models this GPU runs well

Other GPUs that run Mistral Small 3.2 24B

FAQ

How much VRAM does Mistral Small 3.2 24B need?
At Q4_K_M with 8K context: 14.7GB weights + 1.3GB KV cache + 1.5GB runtime overhead = 17.5GB total. RTX 5090 offers 32.0GB usable VRAM, so the verdict is: Comfortable.
What is the best quantization for Mistral Small 3.2 24B on RTX 5090?
Q8_0 — the highest tier that still fits within 32.0GB usable VRAM at 8K context. Lower tiers (Q3/Q2) fit too but cost noticeable quality.
How fast does Mistral Small 3.2 24B run on RTX 5090?
About 53 tokens/s at Q8_0 (theoretical estimate, ±30% in real-world use).

Check another combination in the GPU Checker →

Data verified 2026-09-01