Briefing

Intel Optane Persistent Memory Build Runs 1 Trillion-Parameter LLM at ~4 Tokens/sec

ai-dev
by /u/APFrisco · Llama

Benchmark token throughput on a similar Optane PMem build to validate performance claims.

What to do now

Benchmark token throughput on a similar Optane PMem build to validate performance claims.

Summary

A Reddit user demonstrates a custom build that runs the 1 trillion‑parameter Kimi K2.5 model locally at roughly 4 tokens per second, using 768 GB of Intel Optane Persistent Memory (PMem) in Memory Mode. The PMem acts as RAM with the system’s 192 GB of DDR4 ECC as a cache, allowing the bulk of the model’s sparse expert weights to reside on PMem while the 12 GB RTX 3060 GPU handles the dense layers via llama.cpp’s override‑tensor flag. The user also shows that llama.cpp’s ngl auto and cmoe flags can automatically place tensors, achieving similar performance.

The build includes an Intel Xeon Gold 6246 CPU, TYAN S5630GMRE‑CGN motherboard, ASUS Dual GeForce RTX 3060 OC 12GB GPU, six 32 GB DDR4 ECC sticks, six 128 GB Optane DCPMM modules, a 2 TB WD SN850X NVMe SSD, an ASRock Steel Legend 850W PSU, and a Silverstone Grandia Series case. Intel has discontinued the Optane line, making the modules available secondhand at a fraction of equivalent DRAM cost. The post highlights the potential of memory tiering for large‑scale LLM inference on modest hardware budgets.

Key changes

  • 1 trillion‑parameter Kimi K2.5 runs locally at ~4 tokens/sec using 768 GB Optane PMem in Memory Mode
  • Optane PMem serves as RAM with 192 GB DDR4 ECC as cache
  • Hybrid GPU/CPU inference via llama.cpp’s override‑tensor flag fits 12 GB RTX 3060
  • Alternative llama.cpp flags ngl auto and cmoe also achieve similar placement
  • Build uses Intel Xeon Gold 6246, RTX 3060 12GB, six 32 GB DDR4 ECC, six 128 GB Optane DCPMM
  • Intel discontinued Optane, making modules cheaper secondhand
  • WD SN850X 2 TB NVMe SSD provides fast storage
  • Build demonstrates large‑model inference on limited budget

Affects

none

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting