Briefing

LFM2.5 230M In-Browser Inference at 1,400 Tokens/s with Custom WebGPU Kernels

ai-dev
by /u/xenovatech ·

Run LFM2.5 230M in‑browser using the provided WebGPU kernels to evaluate performance.

What to do now

Run LFM2.5 230M in‑browser using the provided WebGPU kernels to evaluate performance.

Summary

The demo shows the LiquidAI/LFM2.5‑230M model running entirely in the browser using custom WebGPU kernels written by the now‑shut‑down Fable 5 team and Opus 4.8. The video was recorded on an M4 Max, and the model is available on Hugging Face as a GGUF file. The demo is hosted on Hugging Face Spaces under webml‑community/lfm2‑webgpu‑kernels. The model achieves roughly 1,400 tokens per second, demonstrating that small LLMs can run efficiently on consumer hardware with WebGPU acceleration. The post highlights the use of custom kernels to bypass the limitations of standard WebGPU APIs. The demonstration underscores the feasibility of in‑browser inference for lightweight models. The author encourages others to try the demo and report performance. The post provides links to the model and the live demo.

Key changes

  • LFM2.5‑230M runs entirely in-browser with custom WebGPU kernels
  • Kernels written by Fable 5 and Opus 4.8
  • Demo recorded on M4 Max
  • Model available on Hugging Face as GGUF
  • Demo hosted at webml‑community/lfm2‑webgpu‑kernels
  • Achieves ~1,400 tokens per second

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting