Briefing

MTPLX: 2.24× Faster TPS on Apple Silicon Using Native MTP Heads

ai-dev
by /u/YoussofAl · Anthropic OpenAI

Benchmark Qwen 3.6‑27B on an M5 Max to confirm the reported 2.24× speedup and verify accuracy at depth 3.

What to do now

Run "mtplx start wizard" to download Qwen 3.6‑27B and benchmark on your MacBook Pro M5 Max to validate the reported speedup and accuracy.

Summary

MTPLX is a new inference engine that leverages a model’s built‑in MTP heads as speculative drafters, delivering up to 2.25× faster token‑per‑second (TPS) on Apple Silicon without extra memory or an external drafter. The tool works on any MTP‑enabled model, supports mathematically exact temperature sampling with rejection sampling, and offers a fully‑featured CLI that includes a wizard, model download, depth detection, API server, browser and terminal chat, benchmarking suite, diagnostics, fan control, and a 562‑test suite.

On a MacBook Pro M5 Max, MTPLX accelerated Qwen 3.6‑27B 4‑bit MLX from 28 tok/s to 63 tok/s at temperature 0.6, top_p 0.95, top_k 20. The optimal draft depth was found to be D3, balancing acceptance rate and verify time. The engine runs on a patched MLX fork with custom Metal kernels—Innovation‑tape GDN capture, GraphBank, Draft‑only requantised LM head, and Small‑M verify qmv—providing significant speedups in the speculative cycle.

MTPLX also includes a full serving stack: OpenAI‑compatible /v1/chat/completions and /v1/completions, Anthropic‑compatible /v1/messages, a browser chat UI, terminal chat, session bank, and health metrics. The project demonstrates how careful kernel tuning and state‑rollback mechanisms can unlock high‑quality, high‑speed inference on consumer hardware.

Key changes

  • MTPLX uses native MTP heads as speculative drafters, boosting TPS by up to 2.25×
  • Supports any MTP‑enabled model with no external drafter or extra memory
  • Employs exact temperature sampling with rejection sampling, adjustable temperatures
  • Custom Metal kernels: Innovation‑tape GDN capture, GraphBank, Draft‑only requantised LM head, Small‑M verify qmv
  • Full CLI and serving stack with OpenAI/Anthropic‑compatible APIs, browser and terminal chat
  • Qwen 3.6‑27B 4‑bit MLX achieved 63 tok/s on M5 Max at temp 0.6, top_p 0.95, top_k 20 (depth 3 optimal)

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting