GPUFits

Can RTX 4090 run Mistral Small 3.2 24B?

✅ Comfortable
Mistral Small 3.2 24B @ Q4_K_M · 8K context
Needed
17.5 GB
Usable
24.0 GB

VRAM breakdown

Weights (Q4_K_M)14.7 GB
KV cache (8K context)1.3 GB
Runtime overhead1.5 GB
Total17.5 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

Q6_K

Estimated generation speed

~38 tok/s @ Q6_K

Theoretical estimates based on published architecture data and measured GGUF sizes; real-world speed varies ±30%.

Cheaper GPUs that run it

Smaller models this GPU runs well

Other GPUs that run Mistral Small 3.2 24B

FAQ

How much VRAM does Mistral Small 3.2 24B need?
At Q4_K_M with 8K context: 14.7GB weights + 1.3GB KV cache + 1.5GB runtime overhead = 17.5GB total. RTX 4090 offers 24.0GB usable VRAM, so the verdict is: Comfortable.
What is the best quantization for Mistral Small 3.2 24B on RTX 4090?
Q6_K — the highest tier that still fits within 24.0GB usable VRAM at 8K context. Lower tiers (Q3/Q2) fit too but cost noticeable quality.
How fast does Mistral Small 3.2 24B run on RTX 4090?
About 38 tokens/s at Q6_K (theoretical estimate, ±30% in real-world use).

Check another combination in the GPU Checker →

Data verified 2026-09-01