PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!
Test the new ternary Qwen3.6 27B on your workloads to evaluate memory savings and performance gains.
Test the new ternary Qwen3.6 27B on your workloads to evaluate memory savings and performance gains.
Summary
PrismML released Bonsai 27B, a ternary Qwen3.6 27B model that runs on 10GB of memory with near fp16 performance. It uses BitNet/Ternary quantization and is available as a GGUF on Hugging Face. The model is 27B parameters, 10GB memory footprint, 32K context on M4 Pro via llama.cpp. It outperforms 2‑bit quant and is more intelligent than Q4_K_XL. It supports multi‑modal input and a 256K context window. The release includes a demo in the OpenComputer harness that builds an interactive HTML report. The article also mentions upcoming dFlash and unclear MTP support. The model is 100% more intelligent than a comparable 2‑bit quant of Qwen3.6 27B. It uses a ternary GGUF via a llama.cpp fork. The release is early access from PrismML team and includes links to whitepaper, blog, HF collection, llama.cpp fork, MLX fork.
Key changes
- Ternary Qwen3.6 27B runs on 10GB memory
- Near fp16 performance claim refers to benchmarks
- Multi‑modal input and 256K context window
- 32K context on M4 Pro via llama.cpp
- 100% more intelligent than 2‑bit quant of Qwen3.6 27B
- Early access release with whitepaper, blog, HF, llama.cpp fork, MLX fork
- Upcoming dFlash and uncertain MTP support
- Demo in OpenComputer harness builds interactive HTML report