GPUFits

Can Mac Studio M3 Ultra (96GB) run Mistral Small 3.2 24B?

✅ Comfortable
Mistral Small 3.2 24B @ Q4_K_M · 8K context
Needed
17.5 GB
Usable
72.0 GB

VRAM breakdown

Weights (Q4_K_M)14.7 GB
KV cache (8K context)1.3 GB
Runtime overhead1.5 GB
Total17.5 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

FP16

Estimated generation speed

~13 tok/s @ FP16

Theoretical estimates based on published architecture data and measured GGUF sizes; real-world speed varies ±30%.

Cheaper GPUs that run it

Smaller models this GPU runs well

Other GPUs that run Mistral Small 3.2 24B

FAQ

How much VRAM does Mistral Small 3.2 24B need?
At Q4_K_M with 8K context: 14.7GB weights + 1.3GB KV cache + 1.5GB runtime overhead = 17.5GB total. Mac Studio M3 Ultra (96GB) offers 72.0GB usable VRAM, so the verdict is: Comfortable.
What is the best quantization for Mistral Small 3.2 24B on Mac Studio M3 Ultra (96GB)?
FP16 — the highest tier that still fits within 72.0GB usable VRAM at 8K context. Lower tiers (Q3/Q2) fit too but cost noticeable quality.
How fast does Mistral Small 3.2 24B run on Mac Studio M3 Ultra (96GB)?
About 13 tokens/s at FP16 (theoretical estimate, ±30% in real-world use).

Check another combination in the GPU Checker →

Data verified 2026-09-01