SubQ Claims 12M Context Window, Production Model 1M-Preview – Claims vs Reality
Verify SubQ’s real context window and benchmark performance before using it in production.
Verify SubQ’s real context window and benchmark performance before deployment.
Summary
SubQ’s launch thread advertises a 12‑million token context window, yet the production model is labeled SubQ 1M‑Preview, a 1‑million token model. The blog claims 12M is a research result, but the production version differs, raising doubts about marketing consistency.
RULER is reported at 128K tokens, far below the threshold where sparse attention truly benefits. MRCR v2 research scores 83 on a 1M benchmark, but the production model drops to 65.9, lower than Opus 4.6 (78.3) and GPT‑5.5 (74). The homepage lists Opus 4.7 at 32.2, while the blog cites Opus 4.6 at 78.3, showing selective comparison.
Pricing figures also conflict: the homepage claims a 1/5 cost, while the launch thread indicates less than 5% of Opus. The advertised 52× speed over FlashAttention is a kernel‑level metric, not end‑to‑end inference. Sparse attention’s known failure mode—fast until the router prunes a connection—mirrors the MRCR drop pattern. No peer‑reviewed technical report has been released, so the claims remain unverified.
Key changes
- SubQ claims 12M context window but production model is 1M-Preview
- RULER reported at 128K, below sparse attention threshold
- MRCR v2 research 1M: 83, production 65.9, below Opus 4.6 (78.3) and GPT‑5.5 (74)
- Homepage pricing 1/5 cost vs launch thread <5% cost discrepancy
- 52× faster than FlashAttention is kernel‑level, not end‑to‑end
- Sparse attention fails when router prunes a connection
- No peer‑reviewed technical report released yet