Briefing

GGUF quantized Gemma‑4‑26B‑A4B‑NVFP4 released with Docker image for llama.cpp

ai-dev
by /u/catlilface69 · Llama

Pull the Docker image catlilface/llama.cpp:gemma4_26b_nvfp4 and test the GGUF quantized Gemma‑4‑26B‑A4B‑NVFP4 model, noting CPU offloading performance issues.

What to do now

Pull the Docker image catlilface/llama.cpp:gemma4_26b_nvfp4 and test the Gemma‑4‑26B‑A4B‑NVFP4 GGUF model, reporting any performance issues.

Summary

A new GGUF version of the nvidia/Gemma‑4‑26B‑A4B‑NVFP4 model has been uploaded to Hugging Face, but it cannot be run on the main branch of llama.cpp. To address this, the author released a Docker image named catlilface/llama.cpp:gemma4_26b_nvfp4 that bundles the necessary runtime. The quantization was made possible by contributions from ynankani to llama.cpp, enabling the A4B‑NVFP4 format. The author notes that CPU offloading still suffers from performance issues, and only a single 5070Ti GPU was used for testing. Users are invited to test the model and provide feedback on its performance, especially regarding CPU offloading. The Hugging Face repository for the GGUF file can be found at https://huggingface.co/catlilface/Gemma-4-26B-A4B-NVFP4-GGUF. The Docker image can be pulled from the same repository. The author thanks ynankani for the work that made this quantization possible.

Key changes

  • GGUF version of nvidia/Gemma‑4‑26B‑A4B‑NVFP4 released.
  • Not runnable on main branch of llama.cpp; Docker image catlilface/llama.cpp:gemma4_26b_nvfp4 provided.
  • Quantization enabled by ynankani’s contribution to llama.cpp.
  • CPU offloading performance issues remain.
  • HF repo link: https://huggingface.co/catlilface/Gemma-4-26B-A4B-NVFP4-GGUF.
  • Docker image tag catlilface/llama.cpp:gemma4_26b_nvfp4.
  • Feedback requested from users.
  • Only tested on 5070Ti GPU.

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting