Briefing

GB10 Solution Atlas Open‑Source Inference Engine for Qwen Models

ai-dev
by /u/Live-Possession-6726 · Claude Anthropic OpenAI

Deploy Atlas via Docker to serve Qwen3.6-35B-A3B-FP8 with speculative decoding and prefix caching.

What to do now

Deploy Atlas on your GPU cluster, test Qwen3.5‑35B performance, and contribute to the repo.

Summary

The GB10 Solution Atlas has been released as an open‑source inference engine written in pure Rust and CUDA, with a 2.5 GB image and a cold start under two minutes. It replaces the generic Python stack with hand‑tuned CUDA kernels for Blackwell SM120/121, covering attention, MoE, GDN, and Mamba‑2, and supports native NVFP4 and FP8 tensor‑core execution. Multi‑Token Prediction (MTP) speculative decoding is enabled, boosting throughput by up to three times. The engine exposes an OpenAI‑style and Anthropic API on a single port, compatible with Claude Code, Cline, OpenCode, and Open WebUI. Performance benchmarks on a DGX Spark show Qwen3.5‑35B reaching 130 tok/s peak and 111 tok/s sustained, a 3.0–3.3× improvement over vLLM. A Docker command is provided to serve Qwen3.6‑35B‑A3B‑FP8 with speculative decoding and prefix caching. Future plans include a Strix Halo port with Spectral Compute, AMD hardware support, and an RTX 6000 Pro Blackwell release.

Key changes

  • Atlas released as open‑source, pure Rust + CUDA, 2.5 GB image, <2 min cold start
  • Hand‑tuned CUDA kernels for Blackwell SM120/121 covering attention, MoE, GDN, Mamba‑2
  • Native NVFP4 + FP8 tensor‑core support
  • Multi‑Token Prediction (MTP) speculative decoding up to 3× throughput
  • OpenAI + Anthropic API on same port, compatible with Claude Code, Cline, OpenCode, Open WebUI
  • Qwen3.5‑35B peak 130 tok/s, sustained 111 tok/s, 3.0–3.3× faster than vLLM
  • Docker run command: avarok/atlas‑gb10:latest serve Qwen/Qwen3.6‑35B‑A3B‑FP8 …
  • Roadmap: Strix Halo port with Spectral Compute, AMD hardware, RTX 6000 Pro Blackwell

Affects

enterprise internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting