If you had a 384gb 4x blackwell what model would you choose?
Benchmark DeepSeek V4 flash on your 4x RTX PRO 6000s or 8 Blackwells to determine suitability for 20 concurrent users.
Benchmark DeepSeek V4 flash on your 4x RTX PRO 6000s or 8 Blackwells to determine suitability for 20 concurrent users.
Summary
A company is planning to host its own local LLM for internal policy and data management tasks, expecting 2‑3 super users and up to 20 concurrent users. The team currently has a 64GB Mac Studio but is speculating on scaling to 8 Blackwell GPUs or 4x RTX PRO 6000s. They are leaning toward DeepSeek V4 flash for its performance and cost‑efficiency. The post invites community input on the best model fit for their use case. The company also mentions potential expansion to 8 Blackwells and the need to evaluate compute requirements. The discussion includes considerations of concurrency, memory, and inference speed for internal workflows.
Key changes
- Planning to host local LLM for internal policy/data management
- Expecting 2‑3 super users and up to 20 concurrent users
- Current hardware: 64GB Mac Studio
- Considering 8 Blackwell GPUs or 4x RTX PRO 6000s
- Leaning toward DeepSeek V4 flash
- Seeking community input on best model fit
- Potential expansion to 8 Blackwells
- Need to evaluate compute requirements, concurrency, memory, inference speed