Briefing

PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!

ai-dev
by /u/BoogerheadCult · Llama

Test the new ternary Qwen3.6 27B on your workloads to evaluate memory savings and performance gains.

What to do now

Test the new ternary Qwen3.6 27B on your workloads to evaluate memory savings and performance gains.

Summary

PrismML released Bonsai 27B, a ternary Qwen3.6 27B model that runs on 10GB of memory with near fp16 performance. It uses BitNet/Ternary quantization and is available as a GGUF on Hugging Face. The model is 27B parameters, 10GB memory footprint, 32K context on M4 Pro via llama.cpp. It outperforms 2‑bit quant and is more intelligent than Q4_K_XL. It supports multi‑modal input and a 256K context window. The release includes a demo in the OpenComputer harness that builds an interactive HTML report. The article also mentions upcoming dFlash and unclear MTP support. The model is 100% more intelligent than a comparable 2‑bit quant of Qwen3.6 27B. It uses a ternary GGUF via a llama.cpp fork. The release is early access from PrismML team and includes links to whitepaper, blog, HF collection, llama.cpp fork, MLX fork.

Key changes

  • Ternary Qwen3.6 27B runs on 10GB memory
  • Near fp16 performance claim refers to benchmarks
  • Multi‑modal input and 256K context window
  • 32K context on M4 Pro via llama.cpp
  • 100% more intelligent than 2‑bit quant of Qwen3.6 27B
  • Early access release with whitepaper, blog, HF, llama.cpp fork, MLX fork
  • Upcoming dFlash and uncertain MTP support
  • Demo in OpenComputer harness builds interactive HTML report

Affects

internal

Source angles · 2 perspectives

Reddit r/LocalLLaMA
Independent angle

PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB

Open
Reddit r/LocalLLaMA
Independent angle

PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!

Open

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting