Local LLM hardware tools
Can your GPU run that LLM?
Pick a model, pick a GPU — get a clear verdict: fits comfortably, tight, needs multiple cards, or not feasible. With VRAM breakdowns, speed estimates, and cost comparisons.
Model & GPU data verified 2026-09-01
14
open-weight models
12
GPUs & Macs
162
answer pages
13
in-depth guides
Stop guessing. Start knowing.
Every answer is computed from verified model architectures, measured GGUF file sizes, and official GPU specs — not forum folklore.
Three free tools. No signup, no email.
VRAM Calculator
Exact VRAM for any model at any quantization and context length — weights, KV cache, overhead.
Open tool →
GPU Compatibility Checker
Pick a GPU, get every model it can run — with multi-GPU options when one card isn't enough.
Open tool →
Local vs API Cost
When does buying a GPU beat paying per token? Amortized hardware vs live API prices.
Open tool →
New to local LLMs? Start here
A 6-step beginner path: from why models eat VRAM to running your first local model — each step ends with a small hands-on task. No prior knowledge needed.
Straight answers to what everyone's searching
Each answer page breaks down VRAM, recommended quantization, speed, and alternatives.
Can RTX 3060 12GB run Llama 3.1 8B?
✅ Comfortable Full answer →
Can RTX 3090 run Qwen3 32B?
⚠️ Tight Full answer →
Can RTX 5090 run Mistral Small 3.2 24B?
✅ Comfortable Full answer →
Can Mac Studio M3 Ultra (96GB) run gpt-oss-120b?
🔀 Needs multi-GPU Full answer →
Can RTX 4090 run Llama 3.3 70B?
❌ Not feasible Full answer →
Can RTX 5090 run DeepSeek-R1 671B?
❌ Not feasible Full answer →
Go deeper
Long-form guides built on the same verified data.
Best GPU for Local LLMs in 2026: Every Budget Tier
The 2026 buyer's guide for local LLM hardware: best used value (RTX 3090), best new cards (RTX 5090), best big-memory option (Mac Studio), and what to skip.
verified 2026-08-04
Mac vs GPU for LLM Inference: Unified Memory Explained
Apple Silicon unified memory vs NVIDIA VRAM: real bandwidth numbers (273/546/819 GB/s), the 75% usable-memory rule, worked examples, and who should buy which.
verified 2026-08-04
Quantization Explained: Q4 vs Q8 and What You Actually Lose
What GGUF quantization costs: Q8_0 vs Q6_K vs Q5_K_M vs Q4_K_M vs Q3/Q2 — measured file sizes, VRAM and speed math, per-scenario picks, when to go below Q4.
verified 2026-08-04
Every number has a source and a date
Model specs from official Hugging Face configs. Quantization sizes measured from real GGUF files. GPU specs from NVIDIA, AMD and Apple. Re-verified monthly — and every change is logged.