Models / Qwen3.8 Flash-Next
Qwen3.8 Flash-Next VRAM Requirements & GPU Pairing
Qwen3.8-Flash-Next is the September 2026 preview of the Qwen4 architecture: 180B total parameters (including a 51.2B n-gram lookup table and MTP layers) with only ~6B active per token (512 routed experts, top-10 + 1 shared). Of its 48 layers, 36 are Gated DeltaNet linear attention and just 12 are QSA full attention (24 Q heads / 2 KV heads), giving it a native 256K context. Note the license: Qwen Community License 1.0, not Apache-2.0. Q4_K_M weighs ~110.3GB, ~112.6GB total — a 96GB Mac (~72GB usable) cannot hold it, and every GPU in our database fails at Q4.
Realistic paths: a 128GB M5 Max (~96GB usable) runs Q2_K (~73.6GB) comfortably and Q3_K_M (~92.3GB) tight; only the 256GB M5 Ultra runs Q4 comfortably. At 6B active, whatever fits is fast — an M5 Max at 614GB/s theoretically reaches ~125 tok/s, and the n-gram table can be offloaded to cut pressure further. Unsloth GGUFs, Ollama support and Mac deployment write-ups all landed within September — this is the largest new model consumer hardware can almost reach.
Architecture Specs
| Total parameters | 180B |
| Active parameters(MoE: VRAM follows total params, speed follows active params) | 6B |
| Layers | 48 |
| KV heads | 2 |
| Head dim | 256 |
| Native context | 256K |
| License | qwen-community |
| KV bytes per token (fp16) | 96.0 KB |
Hybrid architecture: only 12 of 48 layers are full attention; the 36 Gated DeltaNet layers' KV does not grow with context. The KV column above conservatively applies the standard GQA formula to all 48 layers — real KV is roughly 1/4 of the shown value.
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 191.5 GB | 193.8 GB |
| Q6_K | 147.8 GB | 150.1 GB |
| Q4_K_M | 110.3 GB | 112.6 GB |
| Q3_K_M | 90.0 GB | 92.3 GB |
| Q2_K | 71.3 GB | 73.6 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Not feasible | — | — | Try it → |
| RTX 3090 | 24.0 GB | Not feasible | — | — | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Not feasible | — | — | Try it → |
| RTX 4090 | 24.0 GB | Not feasible | — | — | Try it → |
| RTX 5090 | 32.0 GB | Not feasible | needs 4 cards | — | Try it → |
| RTX A6000 | 48.0 GB | Not feasible | needs 3 cards | — | Try it → |
| A100 80GB | 80.0 GB | Not feasible | Q2_K | ≈643 tok/s | Try it → |
| H100 80GB | 80.0 GB | Not feasible | Q2_K | ≈1057 tok/s | Try it → |
| RX 7900 XTX | 24.0 GB | Not feasible | — | — | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Not feasible | needs 4 cards | — | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Not feasible | needs 3 cards | — | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Not feasible | needs 2 cards | — | Try it → |
| Mac Studio M5 Ultra (96GB) | 72.0 GB | Not feasible | needs 2 cards | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- Why does a 180B model run at 6B speed?
- Sparse MoE activation: only 10 of 512 routed experts plus 1 shared expert fire per token — ~6B active. Generation speed depends on the weights read per token (~3.7GB at Q4_K_M), not the total parameter count.
- Can a 96GB Mac Studio run Qwen3.8-Flash-Next?
- Q4_K_M (~112.6GB) is impossible; even the smallest Q2_K (~73.6GB) misses 72GB usable by 1.6GB, landing in 'needs multi-GPU' territory. The single-machine answers are a 128GB M5 Max (Q2 comfortable / Q3 tight) or a 256GB M5 Ultra (Q4 comfortable).
- What should I check before commercial use?
- It ships under the Qwen Community License 1.0 (not Apache-2.0) — read the commercial and redistribution terms. Also, the 51.2B n-gram table can be offloaded: slightly lower quality, much lower memory pressure.
Related guides
- MoE Model Hardware Requirements: Total vs Active Parameters
- The VRAM Cost of Long Context (and How to Right-Size It)
- Mac vs GPU for LLM Inference: Unified Memory Explained
Try Qwen3.8 Flash-Next in the GPU compatibility checker →
Data verified 2026-10-01