Briefing

OpenAI GPT‑5.5, Codex Update, and Open‑Weight Model Landscape – 4/29/2026

ai-dev
Claude OpenAI DeepSeek

Review GPT‑5.5’s improved cyber task performance and consider leveraging its 60 % lower cost on high‑value workflows.

What to do now

Review GPT‑5.5’s performance metrics and plan to test it on your security workflows.

Summary

OpenAI’s GPT‑5.5 has entered the top tier for long‑horizon cyber tasks, achieving a 71.4 % pass rate on the UK AI Security Institute’s multi‑step cyber‑attack simulation versus 68.6 % for Claude Mythos Preview, and solving the TLO chain in 2/10 attempts compared to Mythos’ 3/10. The model continues to improve beyond a 100 M‑token inference budget, indicating no obvious saturation yet, and is accompanied by a product‑side security release that adds phishing‑resistant sign‑in and hardened recovery for ChatGPT.

Codex has been re‑positioned as a general computer‑use agent, with a new update that introduces role‑based onboarding, app connections, and workflows across documents, slides, spreadsheets, research, and planning. The update delivers a 20 % faster computer/browser use overall, with a 42 % speed increase reported for the “Computer Use” workflow.

The open‑weight arena saw several notable releases: Qwen3.6 27B offers 262 K context, BF16 weights, and 144 M output tokens, costing roughly 21× more per token than Gemma 4 31B; Hy3‑preview runs 295 B total/21 B active MoE with 256 K context and scores 42 on the Intelligence Index; Grok 4.3 climbs to 53, with 40 % lower input price and 60 % lower output price; Ling 2.6 1T delivers a 34 score but a 92 % hallucination rate. DeepSeek’s V4‑Flash vision work and GUI‑agent focus, along with speculation of >100 T tokens for frontier pre‑training, highlight the scale of current research, while Cursor’s new harness‑centric engineering notes a shift toward runtime, evals, and degradation repair.

These developments collectively signal a move from model‑centric bragging to harness‑centric engineering and a broader emphasis on practical, cost‑efficient AI systems.

Key changes

  • GPT‑5.5 achieves 71.4 % pass rate on cyber eval vs 68.6 % for Mythos
  • GPT‑5.5 solves TLO chain 2/10 vs Mythos 3/10
  • Advanced Account Security adds phishing‑resistant sign‑in to ChatGPT
  • Codex update introduces role‑based onboarding, app connections, 20 % faster computer use
  • Qwen3.6 27B offers 262 K context, BF16 weights, 144 M output tokens, 21× cost
  • Grok 4.3 scores 53, 40 % lower input price, 60 % lower output price
  • Ling 2.6 1T scores 34 with 92 % hallucination rate
  • DeepSeek V4‑Flash focuses on vision‑guided GUI agents

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting