Briefing

MTP Performance on mlx-vlm with Gemma4-26b-a4b

ai-dev
by /u/Hydroskeletal ·

MTP improves code generation speed by 1.53× but hurts JSON output; use MTP only when token acceptance >50%.

What to do now

Enable MTP only for code generation workloads with high acceptance; disable for JSON tasks.

Summary

A user tested MTP on mlx‑vlm with the Gemma4‑26b‑a4b model and measured performance across three workloads: code generation, long‑form prose, and JSON output. Code generation saw a 1.53× speedup (75 tok/s to 114.8 tok/s) with a 66% draft acceptance rate. Long‑form prose experienced a slight slowdown (75 tok/s to 71.1 tok/s) with a 31% acceptance rate. JSON output was significantly slower (51.3 tok/s to 25.6 tok/s) with only an 8% acceptance rate. The results suggest that MTP is beneficial for code generation when token acceptance stays above 50%, but it harms JSON generation and long‑form prose when acceptance drops.

The user ran the tests on an M4 Max Studio with Gemma4‑26b‑a4b, highlighting the trade‑offs of MTP in different scenarios.

Key changes

  • Code generation speedup from 75 tok/s to 114.8 tok/s (1.53×) with 66% acceptance
  • Long‑form prose slowed from 75 tok/s to 71.1 tok/s with 31% acceptance
  • JSON output slowed from 51.3 tok/s to 25.6 tok/s with 8% acceptance
  • MTP benefits code generation when acceptance >50%
  • MTP harms JSON and prose when acceptance drops below 50%
  • Tests performed on M4 Max Studio with Gemma4‑26b‑a4b

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting