Briefing

NVIDIA Releases Nemotron‑TwoTower‑30B‑A3B‑Base‑BF16: Diffusion-Based LLM with 2.42× Faster Generation

ai-dev
by /u/nikhilprasanth ·

Integrate Nemotron‑TwoTower‑30B‑A3B‑Base‑BF16 into your inference pipeline and benchmark against autoregressive baseline.

What to do now

Integrate Nemotron‑TwoTower‑30B‑A3B‑Base‑BF16 into your inference pipeline and benchmark against autoregressive baseline.

Summary

NVIDIA has released Nemotron‑TwoTower‑30B‑A3B‑Base‑BF16, a diffusion‑based language model built on the Nemotron 3 Nano 30B‑A3B backbone. The architecture couples a frozen autoregressive context tower with a diffusion denoiser tower that iteratively fills blocks of tokens in parallel. NVIDIA reports that the default mask‑diffusion setup retains 98.7 % of the autoregressive baseline’s aggregate benchmark quality while delivering a 2.42× increase in wall‑clock generation throughput. The model uses BF16 precision and is available on Hugging Face. The release demonstrates that diffusion can be applied to large‑scale language modeling with minimal loss in quality. The announcement positions NVIDIA as a competitor in the LLM space beyond its GPU hardware. The model is expected to attract researchers interested in efficient generation. The release is the first major non‑AI product launch in the NVIDIA ecosystem for LLMs.

Key changes

  • Nemotron‑TwoTower‑30B‑A3B‑Base‑BF16 is diffusion‑based language model
  • Combines frozen autoregressive context tower with diffusion denoiser tower
  • Retains 98.7 % of autoregressive baseline quality
  • Increases wall‑clock generation throughput by 2.42×
  • Uses BF16 precision
  • Built on Nemotron 3 Nano 30B‑A3B backbone

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting