Briefing

vLLM 0.20.0 Adds TurboQuant 2‑bit KV Cache and DeepSeek V4 MegaMoE Support

ai-dev
DeepSeek

Patch vLLM to v0.20.0 and enable TurboQuant 2‑bit KV cache for 4× capacity to reduce memory usage.

What to do now

Patch vLLM to v0.20.0, enable TurboQuant 2‑bit KV cache, and test DeepSeek V4 MegaMoE on your hardware.

Summary

vLLM released version 0.20.0, adding TurboQuant 2‑bit KV cache that quadruples KV capacity and reduces memory usage. The update re‑enables FA4 for MLA prefill on SM90+ GPUs and introduces a new vLLM IR foundation, while fused RMSNorm delivers a 2.1% end‑to‑end latency improvement. vLLM now supports DeepSeek V4 MegaMoE on Blackwell, Jetson Thor, ROCm, Intel XPU, and simplifies GB200/Grace‑Blackwell setup. Early DeepSeek V4 Pro serving results show B300 can be up to 8× faster than H200 for this workload, and DeepGEMM MegaMoE fuses EP dispatch, EP combine, GEMMs, and SwiGLU into a single mega‑kernel. The release also enhances day‑0 support for several open models, including Poolside’s Laguna XS.2, Ling‑2.6‑flash, and NVIDIA’s Nemotron 3 Nano Omni.

These changes aim to lower inference costs and improve performance across a broad range of hardware, making large‑model serving more accessible.

The focus on memory efficiency and hardware‑agnostic deployment reflects the industry’s push toward scalable, cost‑effective inference.

Key changes

  • vLLM v0.20.0 adds TurboQuant 2‑bit KV cache, providing 4× KV capacity.
  • FA4 is re‑enabled for MLA prefill on SM90+ GPUs.
  • New vLLM IR foundation and fused RMSNorm deliver 2.1% end‑to‑end latency improvement.
  • DeepSeek V4 MegaMoE support added for Blackwell, Jetson Thor, ROCm, Intel XPU, and easier GB200/Grace‑Blackwell setup.
  • DeepSeek V4 Pro serving results show B300 can be up to 8× faster than H200 for this workload.
  • DeepGEMM MegaMoE kernel fuses EP dispatch, EP combine, GEMMs, and SwiGLU into a single mega‑kernel.

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting