Briefing

Evaluating dual RTX 6000 vs Mac Studio M3 Ultra for large model inference

ai-dev
by /u/iVoider ·

Test PP speed on a dual RTX 6000 setup for your target model.

What to do now

Set up a dual RTX 6000 system and run PP benchmarks for your model.

Summary

A user wants to run large models such as GLM 5.1 or Kimi k2.6 and compares a Mac Studio M3 Ultra with 512 GB RAM to a hybrid single‑GPU setup (RTX 6000 or 5090) and a system with an EPYC 9xxxx CPU and 12‑channel DDR5 6400 RAM. Benchmarks show that the Mac Studio’s performance is abysmal beyond a 96 k‑token context size, slightly better than the single‑GPU setup. The user wonders whether adding a second RTX 6000 would boost performance by parallelising dense‑model tensors and how much improvement could be expected.

Key observations: Mac Studio M3 Ultra 512 GB RAM, PP speed drops after 96 k tokens, hybrid single‑GPU benchmark, EPYC 9xxxx + DDR5 6400, dual RTX 6000 may improve PP, magnitude uncertain.

The article highlights the trade‑off between CPU‑heavy and GPU‑heavy inference setups for very large context sizes.

Key changes

  • Mac Studio M3 Ultra 512 GB RAM used for inference
  • PP speed is abysmal after 96 k tokens
  • Hybrid single‑GPU benchmark performed
  • System with EPYC 9xxxx and DDR5 6400 used
  • Dual RTX 6000 may improve PP speed
  • Exact improvement magnitude is unknown

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting