Briefing

Zyphra Unveils ZAYA1-8B: 760M Active Parameter MoE Model Trained on AMD GPUs

ai-dev
by steveharing1 · Claude DeepSeek

Test ZAYA1-8B’s 760M active parameter MoE model on math and coding tasks, and experiment with its Markovian RSA inference to gauge compute scaling benefits.

What to do now

Integrate Zyphra’s forked vLLM into your environment and benchmark ZAYA1-8B on math and coding tasks to evaluate its suitability for client projects.

Summary

Zyphra has released ZAYA1-8B, a mixture‑of‑experts language model that runs with only 760 million active parameters while storing 8.4 billion total parameters. The model was trained end‑to‑end on a 1,024‑node AMD Instinct MI300X cluster using IBM‑built Pensando Pollara interconnect, proving that an AMD‑based stack can produce frontier‑level performance without NVIDIA hardware. On math benchmarks it scores 89.1 on AIME 2026, beating DeepSeek‑R1 and Claude Sonnet 4.5, and its custom attention mechanism keeps reasoning quality high even at the low active‑parameter budget. ZAYA1‑8B also introduces a Markovian RSA inference method that generates multiple parallel reasoning traces and keeps the context window bounded, allowing performance to scale with compute. However, the model underperforms on agentic benchmarks such as BFCL‑V4 and IFBench, indicating it is best suited for math, science, and complex coding tasks rather than general chat or tool‑calling. Local deployment requires Zyphra’s forked vLLM; the standard vLLM install will not work. The model’s weights are available under an Apache 2.0 license on Hugging Face, and Zyphra Cloud offers a serverless endpoint for quick testing.

Key changes

  • ZAYA1-8B is a MoE model with 8.4B total params and 760M active at inference
  • Trained entirely on a 1,024‑node AMD Instinct MI300X cluster, proving AMD stack viability
  • Scores 89.1 on AIME 2026, surpassing DeepSeek‑R1 and Claude Sonnet 4.5 on math benchmarks
  • Custom attention mechanism maintains reasoning quality at low active‑parameter budget
  • Introduces Markovian RSA inference that generates parallel reasoning traces and bounds the context window
  • Requires Zyphra’s forked vLLM for local deployment; standard vLLM incompatible
  • Performance scales with compute budget, RSA boosts outperform baseline on LiveCodeBench

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting