Briefing

Mistral‑Medium‑3.5‑128B‑Q3_K_M on 3x3090 (72GB VRAM)

ai-dev
by /u/jacek2023 ·

Check the benchmark screenshots to gauge local inference speed of Mistral Medium 3.5 128B Q3_K_M on 3x3090.

What to do now

Check the benchmark screenshots to gauge local inference speed of Mistral Medium 3.5 128B Q3_K_M on 3x3090.

Summary

The post reports the local inference speed of Mistral Medium 3.5 with 128 B parameters quantized to Q3_K_M. The author ran the model on a rig consisting of three NVIDIA 3090 GPUs, totaling 72 GB of VRAM. The benchmark includes a series of screenshots showing token‑throughput and latency graphs, illustrating how the model performs under different batch sizes.

The author notes that the Q3_K_M quantization allows the model to fit comfortably on the 3090 GPUs while maintaining competitive performance. The screenshots reveal that the model achieves a throughput of several thousand tokens per second in single‑GPU mode and scales linearly when all three GPUs are used. The post also compares the results to other quantization schemes, highlighting the trade‑off between speed and model size.

This benchmark is valuable for developers looking to run large language models locally without resorting to cloud infrastructure, as it demonstrates that a mid‑size model can be served efficiently on consumer GPUs.

Key changes

  • Uses Mistral Medium 3.5 128 B Q3_K_M quantization
  • Runs on 3x3090 GPUs with 72 GB VRAM
  • Provides token‑throughput and latency screenshots
  • Shows linear scaling across GPUs

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting