Local LLM Inference with RTX 3090 vs Sparks and MiniMax M2.7
Consider using two Sparks with MiniMax M2.7 for 120k context coding; they deliver ~15 tok/s at 100k context and 256 GB VRAM, with idle 50 W each.
Evaluate whether two Sparks with MiniMax M2.7 fit your local inference needs; benchmark token throughput and power usage against your current rig.
Summary
An engineer ran a 4‑GPU RTX 3090 rig with 96 GB VRAM and DDR4 2133 RAM to test the Qwen 3.5‑122B‑A10B AWQ model up to 200 k context for web‑app coding. The rig achieved about 15 tokens per second at 100 k context, but the author wants 120 k context, so they are considering two Sparks paired with MiniMax M2.7. Each Spark has 128 GB VRAM, so two give 256 GB total, leaving headroom for the next MiniMax version and Qwen 3.6‑122B. Idle power for a Spark is ~50 W, compared with ~130 W idle for the 4×3090 rig, while the 3090 rig peaks at ~750 W under full load. Output speed stays roughly constant (~15 tok/s) regardless of context size. Two Sparks are needed to keep prompt processing acceptable, and the extra VRAM allows future model upgrades.
Key changes
- 4×RTX 3090 rig provides 96 GB VRAM and 2133 DDR4 RAM
- Qwen 3.5‑122B‑A10B AWQ tested up to 200 k context
- Two Sparks with MiniMax M2.7 each 128 GB VRAM give 256 GB total
- Idle power 50 W per Spark versus 130 W idle for 4×3090
- Peak load 750 W for 122B model on 3090 rig
- Output tokens/sec ~15 tok/s at 100 k context
- Two Sparks needed to keep prompt processing acceptable
- Future models Qwen 3.6‑122B expected