Side‑by‑Side Comparison of Qwen‑Image, ERNIE, and FLUX.2 Dev on RTX 5090
Benchmark Qwen‑Image, ERNIE, and FLUX.2 Dev on RTX 5090 to choose fastest model.
Use ERNIE‑Image Turbo for quick previews and Qwen‑Image for higher quality when time permits.
Summary
The author benchmarks four open‑source image models on a single NVIDIA RTX 5090: Qwen‑Image‑2512 (BF16), ERNIE‑Image Base (BF16), ERNIE‑Image Turbo (BF16 8‑step DMD‑distilled), and FLUX.2 Dev (NVFP4 mixed). Generation times average 55 s for Qwen‑Image, 43 s for ERNIE‑Base, 5 s for ERNIE‑Turbo, and 16 s for FLUX.2 Dev. Qwen‑Image and FLUX.2 Dev spill heavily into system RAM during inference, filling almost all VRAM and system RAM. ERNIE models fit comfortably in VRAM with headroom, though the CPU still dispatches work during sampling. The author notes that lower‑quantized variants of Qwen‑Image produce artifacts on flat surfaces, while BF16 yields better results.
Hardware details include an RTX 5090 with 32 GB VRAM, an AMD Ryzen 9 9950X3D CPU, and 64 GB DDR5 RAM. The author also mentions having FLUX.2 Klein 9B for fast previews but excluding it from the comparison due to style limitations.
The post provides a practical guide for selecting a model based on speed, memory usage, and output quality for a high‑end GPU setup.
Key changes
- Qwen‑Image‑2512 BF16 takes 55 s per image.
- ERNIE‑Image Base BF16 takes 43 s.
- ERNIE‑Image Turbo BF16 8‑step DMD‑distilled takes 5 s.
- FLUX.2 Dev NVFP4 mixed takes 16 s.
- Qwen‑Image and FLUX.2 spill into system RAM during inference.
- ERNIE models fit comfortably in VRAM with CPU dispatch work.
- Hardware: RTX 5090 32 GB, Ryzen 9 9950X3D, 64 GB DDR5.
- Lower‑quantized Qwen‑Image variants produce artifacts on flat surfaces.