Briefing

OpenAI Unveils GPT‑5.5 with Strong Cybersecurity Performance and Codex Enhancements

security
Claude OpenAI

GPT‑5.5 can locate security vulnerabilities like Claude Mythos; integrate it into automated security testing.

What to do now

Add GPT‑5.5 to your security testing workflow and benchmark against Mythos.

Summary

OpenAI has announced the launch of its latest language model, GPT‑5.5, which demonstrates significant improvements in cybersecurity testing and general-purpose computing. In a benchmark run by the UK AI Security Institute, GPT‑5.5 achieved a 71.4 % pass rate on a simulated cyber‑attack scenario, outpacing Claude Mythos Preview’s 68.6 %. The new model also solved the TLO chain challenge in 2 out of 10 attempts, compared with Mythos’ 3 out of 10, and its performance continues to rise even after exceeding 100 million tokens of inference budget. These results suggest that GPT‑5.5 is becoming more reliable in identifying and mitigating potential vulnerabilities.

Alongside the language model, OpenAI has upgraded Codex, positioning it as a general computer‑use agent. The update introduces role‑based onboarding, app connections, and a 20 % increase in computer‑use speed. Users report an overall speed boost of 42 %, indicating that Codex can now handle a wider range of programming and automation tasks more efficiently. In the same announcement, OpenAI rolled out Advanced Account Security for ChatGPT, featuring phishing‑resistant sign‑in methods and hardened recovery options to protect user accounts from social‑engineering attacks.

The company also highlighted developments in the open‑weight model space. Qwen3.6 27B emerged as a key release, offering a 262 k token context window, BF16 weights, and 144 M output tokens, though it is 21 times more expensive than Gemma 4 31B. Meanwhile, Ling 2.6 1T exhibited a 92 % hallucination rate on the AA‑Omniscience benchmark, underscoring ongoing reliability concerns for large open‑weight models. These findings illustrate a broader trend toward more efficient, versatile AI tools, while also pointing to cost and hallucination as persistent challenges.

Overall, GPT‑5.5 and the refreshed Codex are poised to influence internal tooling and productivity across industries, offering stronger security features and faster computation. The open‑weight models continue to close the gap with closed‑source leaders, but their higher costs and higher hallucination rates remain significant hurdles for widespread adoption.

Key changes

  • GPT‑5.5 is publicly available
  • Comparable vulnerability‑finding capability to Claude Mythos
  • Evaluated by the UK AI Security Institute
  • General availability now
  • No specific CVE detection reported

Affects

internal

Source angles · 3 perspectives

Simon Willison
Independent angle

Our evaluation of OpenAI's GPT-5.5 cyber capabilities

Open
unknown
Independent angle

OpenAI GPT‑5.5 Evaluated for Cybersecurity Vulnerability Detection

Open
unknown
Independent angle

OpenAI GPT‑5.5, Codex Expansion, and Open‑Weight Model Landscape

Open

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting