Briefing

Qwen3.6‑35B MoE Runs Smoothly on 12 GB VRAM

ai-dev
by /u/jwestra ·

Use plain decoding with 32 k context and –ncmoe 20 on a 12 GB GPU for coding tasks.

What to do now

Use plain decoding with 32 k context and –ncmoe 20 on a 12 GB GPU for coding tasks.

Summary

Using an RTX 3060 12 GB GPU, the Qwen3.6‑35B‑A3B‑MTP‑IQ4_XS.gguf MoE model demonstrates that 12 GB of VRAM is a practical sweet spot for this 35 B MoE architecture. The benchmark shows plain decoding at ~914 t/s for pp512 and ~46.8 t/s for tg128, with a best configuration of –ncmoe 18, –t 9, –ctk q8_0, –ctv q8_0. A 32 k context size yields ~88.9 t/s prompt throughput and ~43.4 t/s generation while keeping ~273 MiB VRAM free. MTP speculative decoding provides only a ~2 % speedup, reaching ~47.7 t/s generation with –spec‑draft‑n‑max 2. Optimal plain decoding uses –ncmoe 18 for safety, –ncmoe 17 at the edge, and avoids –ncmoe 16. KV cache experiments show q8 KV is essentially free and preferable, with –ctk q8_0 and –ctv q8_0 delivering the highest throughput. The author recommends plain decoding with 32 k context for coding workloads on 12 GB GPUs.

Key changes

  • 12 GB VRAM supports Qwen3.6‑35B‑MoE with ~914 t/s pp512
  • Best config: –ncmoe 18, –t 9, –ctk q8_0, –ctv q8_0
  • 32 k context yields ~88.9 t/s prompt, ~43.4 t/s generation
  • MTP speculative decoding gives ~47.7 t/s, ~2 % speedup
  • KV cache q8 is free and preferable
  • –ncmoe 18 safe, –ncmoe 17 edge, –ncmoe 16 avoid
  • VRAM free ~273 MiB at best config

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting