Briefing

DigitalOcean Achieves 230 tok/s on DeepSeek V3.2, MiniMax‑M2.5, and Qwen 3.5 397B in Serverless Inference

ai-dev
by Bhaskar Dutt · Bedrock DeepSeek

Benchmark your workloads against DigitalOcean’s new DeepSeek V3.2, MiniMax‑M2.5, and Qwen 3.5 397B to validate the 230 tok/s speed advantage.

What to do now

Benchmark your workloads against these new models to confirm the 230 tok/s speed advantage and assess latency.

Summary

DigitalOcean announced the general availability of DeepSeek V3.2, MiniMax‑M2.5, and Qwen 3.5 397B on its Serverless Inference platform, delivering the fastest output speeds across all providers as measured by Artificial Analysis. DeepSeek V3.2 achieves 230 tokens per second with a sub‑1‑second time‑to‑first‑token for 10 000 input tokens, outperforming AWS Bedrock by 3.9× and matching only Google Vertex in TTFT. The performance gains stem from a multi‑layer optimization strategy that includes NVIDIA Blackwell Ultra HGX B300 GPUs with 288 GB HBM3e, 1.5× NVFP4 compute power, NVFP4 quantization reducing memory footprint by ~1.8×, and a custom vLLM serving stack featuring tensor parallelism, kernel fusion, and programmatic dependent launch. The engineering effort also resolved a 25 % performance hit in virtualized environments through close collaboration with NVIDIA, unlocking the full potential of Blackwell silicon. Benchmark results show DeepSeek V3.2, MiniMax‑M2.5, and Qwen 3.5 397B achieving 230 tok/s, 0.96 s TTFT, and balanced latency‑output speed quadrants, positioning DigitalOcean as a leader in inference speed. These advancements enable enterprises to build agentic, real‑time AI applications with predictable latency and lower operational costs. DigitalOcean invites developers to benchmark their workloads against these new models to validate speed gains.

Key changes

  • General availability of DeepSeek V3.2, MiniMax‑M2.5, and Qwen 3.5 397B on Serverless Inference.
  • DeepSeek V3.2 delivers 230 tokens per second and 0.96 s TTFT for 10 000 input tokens, 3.9× faster than AWS Bedrock.
  • Performance achieved using NVIDIA HGX B300 Blackwell Ultra GPUs with 288 GB HBM3e and 1.5× NVFP4 compute, plus NVFP4 quantization reducing memory footprint by ~1.8×.
  • Custom vLLM stack optimizations: tensor parallelism (TP4/TP8), kernel fusion, and programmatic dependent launch.
  • Resolved 25 % performance hit in virtualized environments through collaboration with NVIDIA.
  • Benchmark results place DigitalOcean in the most favorable latency‑output speed quadrant among 12 providers.

Affects

enterprise

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting