Evaluating dual RTX 6000 vs Mac Studio M3 Ultra for large model inference
Test PP speed on a dual RTX 6000 setup for your target model.
Set up a dual RTX 6000 system and run PP benchmarks for your model.
Summary
A user wants to run large models such as GLM 5.1 or Kimi k2.6 and compares a Mac Studio M3 Ultra with 512 GB RAM to a hybrid single‑GPU setup (RTX 6000 or 5090) and a system with an EPYC 9xxxx CPU and 12‑channel DDR5 6400 RAM. Benchmarks show that the Mac Studio’s performance is abysmal beyond a 96 k‑token context size, slightly better than the single‑GPU setup. The user wonders whether adding a second RTX 6000 would boost performance by parallelising dense‑model tensors and how much improvement could be expected.
Key observations: Mac Studio M3 Ultra 512 GB RAM, PP speed drops after 96 k tokens, hybrid single‑GPU benchmark, EPYC 9xxxx + DDR5 6400, dual RTX 6000 may improve PP, magnitude uncertain.
The article highlights the trade‑off between CPU‑heavy and GPU‑heavy inference setups for very large context sizes.
Key changes
- Mac Studio M3 Ultra 512 GB RAM used for inference
- PP speed is abysmal after 96 k tokens
- Hybrid single‑GPU benchmark performed
- System with EPYC 9xxxx and DDR5 6400 used
- Dual RTX 6000 may improve PP speed
- Exact improvement magnitude is unknown