Lucebox DFlash + PFlash on RX 7900 XTX Achieves 2.24× Speedup Over llama.cpp Baseline
Run Lucebox DFlash on 7900 XTX with DDTree budget 8 to achieve ~2.24× speedup over llama.cpp baseline.
Deploy Lucebox DFlash on your 7900 XTX with DDTree budget 8 to achieve ~2.24× speedup over llama.cpp baseline.
Summary
The Lucebox DFlash + PFlash PR #119 was reproduced on an AMD Radeon RX 7900 XTX (gfx1100) with 24 GiB GDDR6 and 62 GiB DDR5, running ROCm 7.1 on Ubuntu 26.04. The test used Qwen3.6‑27B Q4_K_M (15.65 GiB) plus a Lucebox Q8_0 DFlash drafter (1.84 GiB) in a 10‑prompt HumanEval‑style benchmark with --n‑gen 128 and --fast‑rollback. The baseline llama.cpp HIP AR achieved 28.07 tok/s, while DFlash with chain speculation reached 64.23 tok/s (2.29× speedup). DFlash with DDTree budget 8 achieved 62.75 tok/s (2.24× speedup), and budget 22 yielded 60.94 tok/s (2.17×).
Key findings show that budget 8 is optimal on the 7900 XTX, matching the blog’s 2.23× speedup on Strix Halo, and that standard chain speculation is slightly faster than DDTree for short generations. The 7900 XTX’s ~9× bandwidth advantage over Strix Halo’s LPDDR5X explains its higher absolute speed of 62.75 tok/s. The reproduction steps involve cloning the Lucebox repo, installing rocWMMA headers, building for gfx1100, downloading the GGUF models, and running the benchmark script with DFLASH_BIN and DFLASH_DRAFT environment variables. The results confirm that DFlash on 7900 XTX delivers a 2.24× speedup over the llama.cpp baseline with a DDTree budget of 8.
Key changes
- DFlash chain speculation 64.23 tok/s, 2.29× speedup
- DFlash DDTree budget 8 62.75 tok/s, 2.24× speedup
- DFlash DDTree budget 22 60.94 tok/s, 2.17× speedup
- Budget 8 optimal on 7900 XTX
- Standard chain speculation slightly faster than DDTree for short generations
- 7900 XTX absolute speed 62.75 tok/s vs Strix Halo 26.85 tok/s
- ROCm 7.1 and Ubuntu 26.04 used
- Reproduction steps include cloning Lucebox repo, installing rocWMMA, building, downloading GGUFs