Briefing

Qwen 3.6 27B Quantization Test: Chess Board Generation Performance

ai-dev
by /u/bobaburger · Llama

Use IQ4_XS quantization for Qwen 3.6 27B on 16 GB GPUs, enabling turboquant with -ngl 99 to reach 760 tps pp and 22 tps tg.

What to do now

Use IQ4_XS quantization for Qwen 3.6 27B on 16 GB GPUs, enable turboquant with -ngl 99 to reach 760 tps pp and 22 tps tg.

Summary

An experiment was conducted to compare the quality of different quantizations of the Qwen 3.6 27B model when generating SVG chess boards from a PGN prompt. The test measured board state tracking, correct piece placement, orientation, and highlight of the last move.

BF16 served as the full‑precision baseline and produced accurate boards with a 4.7 KB SVG, a correct orientation, and a subtle dotted line. Q8_0 retained almost all features of BF16 but omitted the line, while Q6_K introduced noticeable placement errors on rank‑5 pawns and used a different font. Q5_K_XL matched Q8_0 in visual quality but produced a larger 7.1 KB SVG.

Lower quants such as Q4_K_XL added board coordinates, IQ4_XS delivered the best trade‑off for a 16 GB GPU, achieving 100 tps pp and 8 tps tg with vanilla llama.cpp and up to 760 tps pp / 22 tps tg when turboquant was enabled. IQ3_XXS and below suffered from orientation and alignment issues, making them unsuitable for production use.

Key changes

  • BF16 baseline produces correct board orientation, piece placement, and a 4.7 KB SVG with a dotted line
  • Q8_0 retains BF16 features but removes the line, keeping a similar 4.7 KB SVG
  • Q6_K shows placement errors on rank‑5 pawns and uses a different font, with a 7.1 KB SVG
  • Q5_K_XL matches Q8_0 visually but outputs a larger 7.1 KB SVG
  • IQ4_XS gives the best trade‑off for 16 GB GPUs, achieving 100 tps pp and 8 tps tg with vanilla llama.cpp and up to 760 tps pp / 22 tps tg when turboquant is enabled
  • Lower quants (IQ3_XXS, Q3_K_XL, Q2_K_XL) suffer from orientation or alignment problems

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting