Briefing

StepFun 3.7 Flash: 196B Total Params, 11B Active, Built‑in ViT

ai-dev
by /u/Everlier · OpenAI DeepSeek

Test StepFun 3.7 Flash locally on 128 GB RAM to evaluate agent workflows.

What to do now

Test StepFun 3.7 Flash locally on 128 GB RAM to evaluate agent workflows.

Summary

StepFun has released its 3.7 Flash model, a multimodal Mixture‑of‑Experts architecture with 196 B total parameters but only 11 B active at inference. The model includes a 1.8 B Vision‑Transformer for image tasks and can run locally on a 128 GB RAM machine. Benchmark results show a SWE‑Bench Pro score of 56.26 %, surpassing DeepSeek V4 Flash (55.6 %) and matching Gemini 3.5 Flash (55.1 %). On DeepSearchQA the model achieves an F1 of 92.82 %, close to GPT‑5.5’s 93.98 %. Human‑like Evaluation (HLE) with tools reaches 47.2 %, a solid figure for a flash‑class model. The release is available through OpenRouter and NVIDIA NIM for those who prefer not to self‑host. The high performance relative to its active parameter count makes it an attractive local option for agent‑centric and coding workflows.

The model’s multimodal capabilities and low‑memory footprint position it as a competitive alternative to larger LLMs, especially for teams with limited GPU resources. Its availability on popular inference platforms lowers the barrier to adoption for developers looking to experiment with advanced LLMs without cloud costs.

Key changes

  • 196 B total params, 11 B active at inference
  • Built‑in 1.8 B Vision‑Transformer for image tasks
  • SWE‑Bench Pro score of 56.26 %
  • DeepSearchQA F1 of 92.82 %
  • HLE with tools 47.2 %
  • Runs locally on 128 GB RAM
  • Available via OpenRouter and NVIDIA NIM
  • High performance relative to active parameter count

Affects

none

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting