GLM 5.1, Kimi K2.6, and Qwen 3.5 9B: User‑Driven Insights on Model Performance
Run GLM 5.1 for coding, test Kimi K2.6 if memory allows, switch to Qwen 3.5 9B for multimodal tasks, and monitor the upcoming MTP and novel quantization releases for potential performance gains.
Run GLM 5.1 for coding, test Kimi K2.6 if memory allows, switch to Qwen 3.5 9B for multimodal tasks, and monitor the upcoming MTP and novel quantization releases for potential performance gains.
Summary
The post offers a day‑to‑day review of several large language models, highlighting GLM 5.1 as the most reliable for coding tasks with a trust level of 6, while Kimi K2.6 delivers higher speed but demands 460 GB memory (380 GB when quantized). Minimax 2.7 provides impressive speed for small tasks but is limited in size. Gemma 4 31B remains unstable, with messy mlx support and template bugs, and is not yet ready for production. The author notes that the Qwen 3.6 35B model has been replaced by Qwen 3.5 9B, which offers similar multimodal performance with ~14 GB less memory overhead.
Deepseek 4 Flash and Mimo 2.5 are not yet available in llama.cpp or mlx‑lm, and the author expects their pro versions to be too large for the M3 Ultra. Upcoming releases to watch include Exo, tinygrad, stable Dflash, DDtree, MTP support, and novel quantization formats such as paroquant, JANGTQ, and the pull request 21038. The author also mentions local music generation with Ace Step 1.5, which is almost ready but still lacks voice quality.
Overall, the article serves as a practical guide for selecting models based on memory constraints, speed, and task suitability, while keeping an eye on emerging technologies that could further improve performance.
Key changes
- GLM 5.1 delivers reliable coding support with ~6 trust level; Kimi K2.6 offers higher speed but requires 460 GB memory, 380 GB quantized
- Minimax 2.7 provides good speed for small tasks but limited size
- Gemma 4 31B remains unstable with messy mlx support and template bugs
- Qwen 3.6 35B replaced by Qwen 3.5 9B, offering similar multimodal performance with ~14 GB less memory
- Deepseek 4 Flash and Mimo 2.5 not yet available in llama.cpp or mlx‑lm
- Upcoming releases to watch: Exo, tinygrad, stable Dflash, DDtree, MTP support, novel quantization formats (paroquant, JANGTQ, PR 21038)
- Local music generation with Ace Step 1.5 is almost ready but voice quality still lacking