GPUFits

Can Mac Studio M3 Ultra (96GB) run gpt-oss-120b?

🔀 Needs multi-GPU
gpt-oss-120b @ Q4_K_M · 8K context
Needed
73.6 GB
Usable
72.0 GB

VRAM breakdown

Weights (Q4_K_M)71.5 GB
KV cache (8K context)0.6 GB
Runtime overhead1.5 GB
Total73.6 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

Q3_K_M

Estimated generation speed

~241 tok/s @ Q3_K_M

Theoretical estimates based on published architecture data and measured GGUF sizes; real-world speed varies ±30%.

Smaller models this GPU runs well

FAQ

How much VRAM does gpt-oss-120b need?
At Q4_K_M with 8K context: 71.5GB weights + 0.6GB KV cache + 1.5GB runtime overhead = 73.6GB total. Mac Studio M3 Ultra (96GB) offers 72.0GB usable VRAM, so the verdict is: Needs multi-GPU.
What is the best quantization for gpt-oss-120b on Mac Studio M3 Ultra (96GB)?
Q3_K_M — the highest tier that still fits within 72.0GB usable VRAM at 8K context. Lower tiers (Q3/Q2) fit too but cost noticeable quality.
How fast does gpt-oss-120b run on Mac Studio M3 Ultra (96GB)?
About 241 tokens/s at Q3_K_M (theoretical estimate, ±30% in real-world use).

Check another combination in the GPU Checker →

Data verified 2026-08-04