GPUFits

Models / Mistral Small 3.2 24B

Mistral Small 3.2 24B VRAM Requirements & GPU Pairing

Mistral Small 3.2 24B is Europe's answer to the 24GB single-card sweet spot: 24B dense parameters, 40 layers, 8 KV heads, Apache-2.0, a native 128K context, and a built-in vision encoder. It targets “the strongest dense model a single 24GB card can hold”: Q4_K_M @8K totals ~17.5GB, comfortable on RTX 3090/4090, and Q6_K (~21.3GB) still fits; 16GB cards must drop to Q3.

Against its peers: versus Gemma 3 27B it is smaller, far cheaper on KV (~1.3GB vs 4.2GB at 8K), and freer to license; versus Qwen3 30B-A3B it is dense — every token reads the full ~14.7GB of weights, so a 4090 theoretically manages ~51 tok/s, roughly a quarter of the MoE rival, though dense models have a better reputation for knowledge consistency. General chat, multilingual work, and light vision tasks all fit. It is the default recommendation for 24GB cards.

Architecture Specs

Total parameters24B
Active parameters 24B
Layers40
KV heads 8
Head dim128
Native context128K
Licenseapache-2.0
KV bytes per token (fp16)160.0 KB

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 25.5 GB 28.4 GB
Q6_K 19.7 GB 22.6 GB
Q4_K_M 14.7 GB 17.5 GB
Q3_K_M 12.0 GB 14.8 GB
Q2_K 9.5 GB 12.4 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Not feasible needs 2 cards Try it →
RTX 3090 24.0 GB Comfortable Q6_K ≈36 tok/s Try it →
RTX 4070 Ti Super 16.0 GB Needs multi-GPU Q3_K_M ≈42 tok/s Try it →
RTX 4090 24.0 GB Comfortable Q6_K ≈38 tok/s Try it →
RTX 5090 32.0 GB Comfortable Q8_0 ≈53 tok/s Try it →
RTX A6000 48.0 GB Comfortable Q8_0 ≈23 tok/s Try it →
A100 80GB 80.0 GB Comfortable FP16 ≈32 tok/s Try it →
H100 80GB 80.0 GB Comfortable FP16 ≈52 tok/s Try it →
RX 7900 XTX 24.0 GB Comfortable Q6_K ≈37 tok/s Try it →
Mac mini M4 Pro (48GB) 36.0 GB Comfortable Q8_0 ≈8 tok/s Try it →
Mac Studio M4 Max (64GB) 48.0 GB Comfortable Q8_0 ≈16 tok/s Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Comfortable FP16 ≈13 tok/s Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

Which quant for Mistral Small 24B on a 24GB GPU?
Q4_K_M (~17.5GB) is the comfortable sweet spot; Q6_K (~22.6GB) fits tightly with ~1.5GB to spare; Q8_0 (~28.4GB) does not fit on one card.
Can a 16GB GPU run Mistral Small 24B?
Not at Q4. Q3_K_M (~14.8GB) barely fits, but quality loss at Q3 is noticeable at this size — the better 16GB answer is gpt-oss-20b (MoE).
Mistral Small 24B or Gemma 3 27B?
Pick Mistral for the Apache-2.0 license, lower VRAM pressure, and long-context headroom. Pick Gemma for Google's English quality and instruction following — accepting a tight fit on 24GB and much heavier KV cache.

Related guides

Try Mistral Small 3.2 24B in the GPU compatibility checker →

Data verified 2026-09-01