GPUFits

Data Updates

Our model, GPU, quantization and API pricing datasets are re-verified on the first weekend of every month. Each entry below lists the date and the source. RSS feed

· Monthly refresh: 2 new models, GPU used-market repricing, API price changes

  • New models: Qwen3.8-27B (Alibaba, Aug 14 — Apache 2.0, hybrid linear/full attention, 262K context; HF config cross-checked via kingy.ai / zenn.dev / NVIDIA developer forums) and Muse Glimmer 30B (Meta, Aug 10 — Apache 2.0, 52 layers, 32Q/2KV GQA, 128K context; config cross-checked via localaimaster / kingy.ai / swfte). Qwen3.8-2.4T-A95B reviewed but excluded: far beyond consumer hardware.
  • New GPUs: none launched in August 2026 (RTX 50 Super / RX 9080 XT / M5 Ultra Mac Studio remain unannounced).
  • GPU used prices (US market snapshot, buysellram 2026-08-07/08 + gpupoet/gpudojo): RTX 3060 12GB $180→$260, RTX 3090 $1,000→$1,100, RTX 4070 Ti Super $550→$760, RTX 4090 $1,100→$2,400 (discontinued; priced against RTX 5090). RTX 5090 note updated: secondary market $3,590–4,304 used / $4,288+ new. RTX A6000 and RX 7900 XTX checked, no confirmed >10% move.
  • API pricing: DeepSeek switched to peak/off-peak pricing on Aug 16 (off-peak V4-Flash $0.22/$0.66, V4-Pro $0.66/$1.98; peak doubles) — api-docs.deepseek.com via codersera/chat-deep.ai. OpenAI (openai.com): GPT-5.6 Terra cut to $2/$12 on Jul 30, GPT-5.6 Luna added at $0.20/$1.20; Sol promo $4/$20 noted. Anthropic cancelled the planned Sep 1 Sonnet 5 increase — $2/$10 is now permanent (Anthropic pricing note, Aug 19). Google: Gemini 3.7 Flash added ($0.75/$3.75 intro through Dec 31). Together: Llama 3.3 70B $0.88→$1.04. Groq / DeepInfra: unchanged.

· Initial dataset verified and published

  • 12 models: specs (params, layers, KV heads, head dim, context) verified against Hugging Face configs; MLA branch confirmed for DeepSeek-R1.
  • 12 GPUs: VRAM, memory bandwidth and MSRP verified against NVIDIA / AMD / Apple official specifications; used-market prices (RTX 3090/4090) sampled from eBay sold listings, July 2026.
  • Quantization presets: bits-per-weight calibrated from measured GGUF file sizes (bartowski, Llama-3.1-8B): FP16 16.0, Q8_0 8.51, Q6_K 6.57, Q5_K_M 5.71, Q4_K_M 4.90, Q3_K_M 4.00, Q2_K 3.17.
  • API pricing: input/output prices per million tokens verified for DeepSeek, Together AI, Groq, DeepInfra, Fireworks, OpenRouter and Novita (7 providers).
  • Apple lineup update: 256GB/512GB Mac Studio M3 Ultra configurations discontinued (DRAM shortage); 96GB is the ceiling, base price now $5,299.