Models / Mistral Small 3.2 24B
Mistral Small 3.2 24B VRAM Requirements & GPU Pairing
Mistral Small 3.2 24B is Europe's answer to the 24GB single-card sweet spot: 24B dense parameters, 40 layers, 8 KV heads, Apache-2.0, a native 128K context, and a built-in vision encoder. It targets “the strongest dense model a single 24GB card can hold”: Q4_K_M @8K totals ~17.5GB, comfortable on RTX 3090/4090, and Q6_K (~21.3GB) still fits; 16GB cards must drop to Q3.
Against its peers: versus Gemma 3 27B it is smaller, far cheaper on KV (~1.3GB vs 4.2GB at 8K), and freer to license; versus Qwen3 30B-A3B it is dense — every token reads the full ~14.7GB of weights, so a 4090 theoretically manages ~51 tok/s, roughly a quarter of the MoE rival, though dense models have a better reputation for knowledge consistency. General chat, multilingual work, and light vision tasks all fit. It is the default recommendation for 24GB cards.
Architecture Specs
| Total parameters | 24B |
| Active parameters | 24B |
| Layers | 40 |
| KV heads | 8 |
| Head dim | 128 |
| Native context | 128K |
| License | apache-2.0 |
| KV bytes per token (fp16) | 160.0 KB |
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 25.5 GB | 28.4 GB |
| Q6_K | 19.7 GB | 22.6 GB |
| Q4_K_M | 14.7 GB | 17.5 GB |
| Q3_K_M | 12.0 GB | 14.8 GB |
| Q2_K | 9.5 GB | 12.4 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Not feasible | needs 2 cards | — | Try it → |
| RTX 3090 | 24.0 GB | Comfortable | Q6_K | ≈36 tok/s | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Needs multi-GPU | Q3_K_M | ≈42 tok/s | Try it → |
| RTX 4090 | 24.0 GB | Comfortable | Q6_K | ≈38 tok/s | Try it → |
| RTX 5090 | 32.0 GB | Comfortable | Q8_0 | ≈53 tok/s | Try it → |
| RTX A6000 | 48.0 GB | Comfortable | Q8_0 | ≈23 tok/s | Try it → |
| A100 80GB | 80.0 GB | Comfortable | FP16 | ≈32 tok/s | Try it → |
| H100 80GB | 80.0 GB | Comfortable | FP16 | ≈52 tok/s | Try it → |
| RX 7900 XTX | 24.0 GB | Comfortable | Q6_K | ≈37 tok/s | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Comfortable | Q8_0 | ≈8 tok/s | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Comfortable | Q8_0 | ≈16 tok/s | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Comfortable | FP16 | ≈13 tok/s | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- Which quant for Mistral Small 24B on a 24GB GPU?
- Q4_K_M (~17.5GB) is the comfortable sweet spot; Q6_K (~22.6GB) fits tightly with ~1.5GB to spare; Q8_0 (~28.4GB) does not fit on one card.
- Can a 16GB GPU run Mistral Small 24B?
- Not at Q4. Q3_K_M (~14.8GB) barely fits, but quality loss at Q3 is noticeable at this size — the better 16GB answer is gpt-oss-20b (MoE).
- Mistral Small 24B or Gemma 3 27B?
- Pick Mistral for the Apache-2.0 license, lower VRAM pressure, and long-context headroom. Pick Gemma for Google's English quality and instruction following — accepting a tight fit on 24GB and much heavier KV cache.
Related guides
- Best GPU for Local LLMs in 2026: Every Budget Tier
- Quantization Explained: Q4 vs Q8 and What You Actually Lose
- KV Cache Explained: Why Long Contexts Eat Your VRAM
Try Mistral Small 3.2 24B in the GPU compatibility checker →
Data verified 2026-09-01