APEX MoE Quantization Update: 30+ Models, New I‑Nano Tier
Add the I‑Nano tier to your quantization scripts, update model references to the new 30+ MoE list, and run long‑context tests on your 30‑50B models.
Add the I‑Nano tier to your quantization scripts, update model references to the new 30+ MoE list, and run long‑context tests on your 30‑50B models.
Summary
APEX, the MoE‑aware mixed‑precision quant strategy, has expanded its collection to over 30 models across major families such as Qwen, Frontier‑size MoEs, Hybrid Mamba/SSM, Gemma 4, and community merges. The new ultra‑compressed I‑Nano tier (IQ2_XXS) reduces mid‑layer routed experts to 2.06 bpw, near‑edge to IQ2_S, edges to Q3_K, and shared experts to Q5_K, making it roughly 20 % smaller than the I‑Mini tier (e.g., Qwen 3.5 35B‑A3B drops from 13 GB to 11 GB). Long‑context performance remains strong: I‑Balanced and I‑Compact keep coherence past 32 k tokens on 30‑50 B MoEs, outperforming uniform Q4_K. Coding quants also perform well; Qwen3.6 35B‑A3B I‑Compact and I‑Mini stay close to F16 on code tasks. The I‑Nano tier requires an imatrix and enables sparse per‑token expert activation, making it viable only on MoE models. The collection now includes new Qwen lineages (122B‑A10B, 397B‑A17B, Claude‑Distilled, Fernflower, TQ), Frontier‑size MoEs (MiniMax‑M2.5, M2.7, Mistral‑Small 4 119B‑2603, NVIDIA Nemotron‑3‑Super 120B‑A12B, GLM‑4.7 Flash, Step‑3.5 Flash, Nemotron‑3‑Nano 30B‑A3B, Omni Reasoning, Holo3 35B‑A3B, Huihui3.5 67B‑A3B), Hybrid Mamba/SSM MoEs (Nemotron‑3‑Nano, Holo3, LFM2 24B‑A2B), Gemma 4 family (gemma‑4 26B‑A4B‑it, Claude Opus distill, heretic, Gemopus‑4 Preview), and community MoE merges (Carnice MoE 35B‑A3B, Carnice‑Qwen3.6, Qwopus MoE 35B‑A3B).
Key changes
- APEX now includes an ultra‑compressed I‑Nano tier (IQ2_XXS) reducing mid‑layer routed experts to 2.06 bpw, near‑edge IQ2_S, edges Q3_K, shared experts Q5_K
- I‑Nano tier is ~20 % smaller than I‑Mini (e.g., Qwen 3.5 35B‑A3B drops from 13 GB to 11 GB)
- Collection expanded to 30+ MoE models across Qwen, Frontier‑size MoEs, Hybrid Mamba/SSM, Gemma 4, and community merges
- Long‑context performance: I‑Balanced and I‑Compact retain coherence beyond 32 k tokens on 30‑50 B MoEs, outperforming uniform Q4_K
- Coding quants: Qwen3.6 35B‑A3B I‑Compact and I‑Mini stay close to F16 on code tasks
- I‑Nano requires an imatrix and enables sparse per‑token expert activation, making it MoE‑only
- New quantization strategy supports sparse per‑token expert activation for efficient deployment
- Links: HuggingFace collection https://huggingface.co/collections/mudler/apex-quants-gguf and GitHub project https://github.com/mudler/apex-quant