Llama405b q4 Performance on AMD Epyc 9374f
Run Llama405b q4 on AMD Epyc 9374f to achieve 1.2k tokens/sec, while newer models reach 30‑100k tokens/sec.
Benchmark your target models on the same hardware to estimate throughput.
Summary
A user reports that Llama405b quantised to q4 runs at 1.2k tokens per second on a 2‑year‑old AMD Epyc 9374f system. The same hardware now achieves 30k‑100k tokens per second when running newer large models such as Kimik2.6, DeepSeekV4Flash, Minimax2.7, Step3.5Flash, and Qwen3.5‑397B, vastly outperforming Llama405b.
The post highlights the dramatic performance gains achievable with newer models and the same hardware, suggesting that local inference of state‑of‑the‑art models is now feasible at a fraction of the cost.
The author encourages readers to experiment with these models and ignore naysayers, noting that the "stupid" experiments are paying off.
Key changes
- Llama405b q4 at 1.2k tokens/sec on older hardware
- Same hardware now runs Kimik2.6, DeepSeekV4Flash, Minimax2.7, Step3.5Flash, Qwen3.5‑397B at 30‑100k tokens/sec
- Demonstrates performance jump
- Suggests cost‑effective local inference