Briefing

Poor Vulkan Performance on Intel Arc 130T for Gemma 4 E4B

ai-dev
by /u/TuskNaPrezydenta2020 · Llama

Use Vulkan on Zen 4 iGPUs for better throughput; consider alternative GPUs if Arrow Lake iGPUs are required.

What to do now

Use Vulkan on Zen 4 iGPUs for better throughput; consider alternative GPUs if Arrow Lake iGPUs are required.

Summary

A developer reports that running llama.cpp with SYCL on an Arrow Lake system yields poor pp/tg performance: 100 tokens/s for pp256 and less than 4 for tg64 using Gemma 4 E4B on an Intel Arc 130T. The developer finds Vulkan easier to set up and achieves better tg speed on Zen 4 iGPUs with Vulkan. The Arrow Lake iGPUs underperform compared to other CPUs. The post asks whether SYCL improves performance on Intel iGPUs or if other solutions are recommended.

This is a performance troubleshooting note for developers using Intel iGPUs with llama.cpp and highlights the current limitations of Vulkan on Arrow Lake.

The information is relevant for those selecting GPU hardware for local LLM inference.

Key changes

  • pp/tg 100 tokens/s for pp256, <4 for tg64 on Arc 130T
  • SYCL attempted but Vulkan easier
  • better tg speed on Zen 4 iGPUs with Vulkan
  • Arrow Lake iGPUs underperform
  • consider alternative GPUs

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting