AI Twitter Recap: Research Benchmarks, Agentic Systems, and Inference Advances
Patch training pipelines to include Soohak's 439 math problems and Medmarks v1.0 benchmarks, and test Perceptron Mk1 on video workloads.
Patch training pipelines to include Soohak's 439 math problems and Medmarks v1.0 benchmarks, and test Perceptron Mk1 on video workloads.
Summary
Soohak released 439 research‑level math problems authored by 64 mathematicians, 38 of whom are faculty, and Medmarks v1.0 expanded its open medical benchmark suite from 20 to 30 benchmarks and from 46 to 61 models, while sentiment grows that old evals are saturating.
DeepMind’s AI Co‑Mathematician achieved 48 % on FrontierMath Tier 4 and Gemini 3.1 Pro jumped from 17.7 % to 31.4 % on CritPt by decomposing tasks into specialized agents, and ProgramBench’s first task was solved by GPT‑5.5 high/xhigh, outperforming Opus 4.7 xhigh.
LightOn’s Agent‑ModernColBERT outperformed Reason‑ModernColBERT by ~10 % on BrowseComp‑Plus with a 149 M‑parameter retriever, while SOAP‑Muon set a new record of 3,150 steps and MuLoCo‑style outer Nesterov SGD improved results, and a Lean4‑to‑TileLang superoptimizer automatically discovered FlashAttention2, FlashNorm, and split‑k matmul, yielding ~1.8× speedup on A100s.
Training‑time tricks such as Nous’s Lighthouse Attention and Prime Intellect’s Renderers reduced pre‑training cost and increased throughput by >3×, and scaling‑law discussions shifted from “20 tokens per parameter” to measuring in bytes.
Inference saw Blackwell racks become the reference for large‑MoE serving, with Perplexity deploying Qwen3 235B on NVIDIA GB200 NVL72, cutting NVLS all‑reduce latency from 586.1 µs on H200 to 313.3 µs on GB200 and MoE prefill combine from 730.1 µs to 438.5 µs, while Modal’s dedicated inference stack and SemiAnalysis’s B200 clustering over RoCEv2 CX‑7 PD disaggregation boosted per‑GPU token throughput by up to 7×.
Qdrant 1.18 added TurboQuant, delivering recall near scalar quantization with 2× less memory, and the Shepherd agent runtime introduced Git‑like version control for agent execution, achieving a 54.7 % success rate on CooperBench.
The Perceptron Mk1 model, launched by Perceptron, offers frontier video and embodied reasoning with native video support up to 2 FPS, temporal grounding, multimodal in‑context learning, and structured spatial outputs, boasting a 32k multimodal context and first‑class outputs like points, boxes, polygons, and clips.
Google DeepMind’s AI‑enabled mouse pointer demos and other multimodal interaction layers also appeared, underscoring a shift toward more interactive AI interfaces.
Key changes
- Soohak introduced 439 research‑level math problems authored by 64 mathematicians
- Medmarks v1.0 expanded benchmarks from 20→30 and models from 46→61
- DeepMind AI Co‑Mathematician reached 48 % on FrontierMath Tier 4
- Gemini 3.1 Pro improved to 31.4 % on CritPt via agent decomposition
- ProgramBench solved first task with GPT‑5.5 high/xhigh, outperforming Opus 4.7 xhigh
- LightOn Agent‑ModernColBERT outperformed Reason‑ModernColBERT by ~10 % on BrowseComp‑Plus
- SOAP‑Muon set record 3,150 steps, and Perceptron Mk1 offers 2 FPS video and 32k multimodal context
- Blackwell racks enabled Qwen3 235B on GB200, cutting latency and boosting throughput