Briefing

GLM 5.2 Unsloth Quant on Consumer Hardware Achieves ~12t/s Inference

ai-dev
by /u/phwlarxoc · Sage Llama

Set up llama.cpp with the given flags to test GLM‑5.2‑UD‑Q5_K_S on dual RTX 5090 for ~12 t/s.

What to do now

Set up llama.cpp with the given flags to test GLM‑5.2‑UD‑Q5_K_S on dual RTX 5090 for ~12 t/s.

Summary

The poster ran the unsloth quant of GLM 5.2, UD‑Q5_K_S, on consumer hardware consisting of a 32‑core Zen5 Threadripper Pro 9975 WX, an Asus WRX90E‑SAGE‑SE PCIe Gen5 motherboard, 512 GB DDR5 ECC RAM @ 4800 MHz, and dual RTX 5090 GPUs. The quant weighs 492 GB and was compiled with llama.cpp using GGML_CUDA, flash‑attn, and several advanced flags. The inference speed achieved was consistently 12 t/s, and the performance was stable regardless of variations in llama.cpp options. The setup demonstrates that large‑scale models can run on high‑end consumer GPUs with proper quantisation. The poster notes that the machine was assembled before the recent RAM price surge, keeping costs lower. The experiment highlights the feasibility of running GLM 5.2 locally without a dedicated data‑center GPU cluster. The post invites others to replicate the setup and share results. The detailed build flags provide a reference for future optimisations.

Key changes

  • GLM‑5.2‑UD‑Q5_K_S quant uses 492 GB of weights
  • Runs on 32‑core Zen5 Threadripper Pro 9975 WX with 512 GB DDR5 ECC RAM @ 4800 MHz
  • Dual RTX 5090 GPUs (CUDA 12.0) used
  • llama.cpp built with GGML_CUDA and flash‑attn enabled
  • Achieves ~12 t/s inference speed
  • Performance stable across various llama.cpp flags

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting