GPUFits

Can RTX 4090 run gpt-oss-120b?

❌ Not feasible
gpt-oss-120b @ Q4_K_M · 8K context
Needed
73.6 GB
Usable
24.0 GB

VRAM breakdown

Weights (Q4_K_M)71.5 GB
KV cache (8K context)0.6 GB
Runtime overhead1.5 GB
Total73.6 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

Does not fit at any quantization

4× RTX 4090 gives 96.0GB usable combined — enough to run it (Comfortable).

Smaller models this GPU runs well

FAQ

How much VRAM does gpt-oss-120b need?
At Q4_K_M with 8K context: 71.5GB weights + 0.6GB KV cache + 1.5GB runtime overhead = 73.6GB total. RTX 4090 offers 24.0GB usable VRAM, so the verdict is: Not feasible.
What is the best quantization for gpt-oss-120b on RTX 4090?
None — even Q2_K (48.4GB) exceeds this GPU's 24.0GB usable VRAM. Use a smaller model, multiple GPUs, or a cloud GPU.

Check another combination in the GPU Checker →

Data verified 2026-08-04