FP8, MXFP8, NVFP4 Quantization Formats: Hardware & Software Support Overview
Enable CUDA 13.0, PyTorch 2.10+cu130, and TorchAO to use MXFP8/NVFP4 on Blackwell GPUs; note that FP8 is only software fallback on RTX 20/30.
Upgrade to CUDA 13.0, install PyTorch 2.10+cu130, and test MXFP8/NVFP4 pipelines on Blackwell GPUs to validate speed and quality.
Summary
This comprehensive guide details the hardware and software support for FP8, MXFP8, and NVFP4 quantization formats across NVIDIA GPU generations. RTX 20/30 (Turing/Ampere) lack native FP8 tensor cores and rely on software fallback, while RTX 40 (Ada Lovelace) offers native FP8 via 4th‑gen Tensor Cores (compute capability 8.9+). The first consumer GPUs with native MXFP8 and NVFP4 are the RTX 50 (Blackwell) series, which also support NVFP4 natively.
Software support has progressed from CUDA 12.6 (stable FP8) to CUDA 13.0 (stable MXFP8/NVFP4). PyTorch 2.1 introduced experimental FP8 dtypes, broadly usable from 2.2 and mature from 2.3+. Stable pip installs for MXFP8/NVFP4 are available with PyTorch 2.10+cu130, released on January 21, 2026, and CUDA 13.0 became stable in that release.
The article also compares quality, VRAM usage, speed, and stability across FP16, FP8, MXFP8, and NVFP4, citing FLUX.1‑Dev benchmarks that show a 1.50x speedup and 1.68x speedup on Blackwell GPUs for NVFP4 versus BF16. It concludes with practical guidance for home users on selecting the right precision for inference workloads.
Key takeaways include the necessity of CUDA 13.0, PyTorch 2.10+cu130, and TorchAO for MXFP8/NVFP4 pipelines, as well as the limited benefit of FP8 on older RTX 20/30 cards.
Key changes
- RTX 40 supports native FP8 via 4th‑gen Tensor Cores (compute 8.9+); RTX 50 adds native MXFP8 and NVFP4; CUDA 13.0 stable for MXFP8/NVFP4; PyTorch 2.10+cu130 provides stable MXFP8/NVFP4 support; FP8 dtypes experimental in 2.1, mature in 2.3+; NVFP4 yields 1.68x speedup on Blackwell vs BF16; MXFP8 offers near‑FP16 quality with lower VRAM; FP8 is software fallback on RTX 20/30 with no speedup; PyTorch 2.10+cu130 requires CUDA 13.0 and driver 570+