Briefing

Qwen 3.6 27B Q8 on four Nvidia RTX A4000 (16GB each) with Llama.cpp and MTP enabled

ai-dev
by /u/Alternative_Ad4267 · Llama

Run Qwen 3.6 27B Q8 on four RTX A4000 GPUs with Llama.cpp and MTP to achieve ~45 tokens/s for reasoning and ~60 tokens/s for coding.

What to do now

Configure your Llama.cpp server with --spec-type mtp, --spec-draft-n-max 4, and run Qwen 3.6 27B Q8 on four RTX A4000 GPUs to benchmark your own throughput.

Summary

The author runs Qwen 3.6 27B Q8 on four Nvidia RTX A4000 GPUs (16 GB each) using Llama.cpp with MTP enabled. The setup uses a Lenovo ThinkStation P3 Tower Gen 2, with each card capped to 125 W to maintain stability. The best performance is achieved with --spec‑draft‑n‑max 4 for MTP, and the author uses Fedora 43 with CUDA drivers. Reasoning throughput reaches ~45 tokens/s, while coding throughput reaches ~60 tokens/s. The author also benchmarks a Qwen 3.6 35B A3B Q8 MoE model, achieving ~90 tokens/s coding and ~80 tokens/s reasoning, though it only fits in layer mode and has lower energy usage. The Qwen 3.6 27B Q8 variant is run as a GGUF with MTP enabled, and the author notes that the MoE model converges slower but is more energy efficient. The author concludes that older 2020 GPUs still deliver respectable performance when properly configured. The setup demonstrates that a heterogeneous cluster can achieve high throughput with modest hardware.

Key changes

  • Uses four RTX A4000 GPUs (16 GB each) capped to 125 W
  • MTP spec-draft-n-max 4 yields best performance
  • Qwen 3.6 27B Q8 variant on GGUF with MTP enabled
  • Reasoning throughput ~45 tokens/s, coding throughput ~60 tokens/s
  • Qwen 3.6 35B A3B Q8 MoE achieves ~90 tokens/s coding, ~80 tokens/s reasoning
  • MoE model fits only in layer mode, uses less energy
  • Uses Fedora 43 with CUDA drivers
  • Tensor split mode with --split-mode tensor

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting