Briefing

Databricks releases GPT‑5.5, achieving 50% accuracy on OfficeQA Pro and 46% error reduction

ai-dev
OpenAI

Migrate agent workflows to GPT‑5.5 via AI Unity Gateway to leverage improved parsing and error reduction.

What to do now

Migrate agent workflows to GPT‑5.5 via AI Unity Gateway to leverage improved parsing and error reduction.

Summary

Databricks announced GPT‑5.5, a new LLM that sets a new state‑of‑the‑art on the OfficeQA Pro benchmark, which tests parsing, retrieval, and grounded reasoning across complex enterprise documents. GPT‑5.5 achieved 50% accuracy, surpassing the 5.4 version, and reduced errors by 46% in agent‑harness workflows. The model shows a step‑function lift in parsing scanned PDFs and legacy files, and it retrieves relevant context more reliably, reducing unnecessary search detours. Databricks now makes GPT‑5.5 available through its AI Unity Gateway, allowing customers to use the model inside workflows built with AgentBricks and the Agent Supervisor API. The release is expected to improve reliability and efficiency for production agent systems that handle long‑context documents. The upgrade also brings improvements in orchestration across multi‑step tasks, better knowledge lift, and more reliable completion of complex workflows without additional supervision. Databricks anticipates widespread adoption of GPT‑5.5 in custom agent workflows, especially where parsing accuracy and error rates are critical.

Key changes

  • GPT‑5.5 achieves 50% accuracy on OfficeQA Pro benchmark
  • 46% error reduction compared to GPT‑5.4 in agent‑harness workflows
  • Step‑function lift in parsing scanned PDFs and legacy documents
  • Improved retrieval of relevant context, reducing unnecessary search detours
  • Available through Databricks AI Unity Gateway for AgentBricks and Agent Supervisor API
  • Enhanced orchestration across multi‑step tasks
  • Better knowledge lift and reliable completion of complex workflows
  • Databricks expects widespread adoption in custom agent workflows

Affects

enterprise

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting