Briefing

Building a 2.3 TB RAM, 400+ vCore Heterogeneous Cluster with Blackwell and Tinygrad

ai-dev
by /u/Street-Buyer-2428 ·

Explore heterogeneous cluster setups for large‑scale inference.

What to do now

Contact the author via PM to discuss collaboration.

Summary

A Reddit user reports assembling a massive heterogeneous AI cluster featuring 2.3 TB of RAM and over 400 virtual cores. The setup uses Blackwell GPUs for pre‑fill and RDMA to a studio mesh for decoding, aiming to be the first of its kind. The user notes that the Tinygrad driver must be integrated to make the cluster functional, and that the remaining step is to connect the hardware to the Blackwell driver. The post highlights the potential for large‑scale inference workloads on mixed‑hardware environments.

The cluster would leverage Blackwell’s high‑bandwidth interconnects for rapid data transfer, while RDMA enables low‑latency communication between nodes. Tinygrad, a lightweight deep‑learning framework, is required to orchestrate the workload across the heterogeneous resources. The author invites collaboration from others with expertise in these domains.

This information is useful for teams exploring advanced AI infrastructure beyond standard cloud offerings, especially those interested in custom hardware acceleration and low‑latency inference pipelines.

Key changes

  • 2.3 TB RAM
  • 400+ vCores
  • Blackwell GPUs for pre‑fill
  • RDMA to studio mesh for decode
  • Tinygrad driver required
  • First heterogeneous cluster
  • Connection to Blackwell driver pending

Affects

none

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting