Safety Helping ChatGPT better recognize context in sensitive conversations
Update ChatGPT safety model to generate safety summaries that capture earlier risk context and use them to trigger more cautious responses in high‑risk conversations.
Update: Deploy the new safety model and monitor safe‑response metrics for high‑risk conversation scenarios.
Summary
OpenAI introduced safety updates that generate safety summaries capturing earlier risk context, enabling ChatGPT to recognize subtle or evolving cues in high‑risk conversations.
The updates improved safe‑response performance by 50% in suicide and self‑harm scenarios and 16% in harm‑to‑others, with GPT‑5.5 Instant showing 52% and 39% improvements respectively.
Safety summaries are short, factual, limited‑time notes used only for serious safety concerns, and were evaluated across 4,000+ cases with a safety relevance score of 4.93/5.
Training involved mental health experts, and the updates did not degrade ordinary conversation quality.
The system now de‑escalates, refuses harmful details, or redirects to safer alternatives when risk emerges over time.
Future work will explore extending these methods to other high‑risk areas such as biology or cyber safety, with safeguards in place.
Key changes
- Introduced safety summaries capturing earlier safety‑relevant context.
- Improved safe‑response performance by 50% in suicide/self‑harm, 16% in harm‑to‑others.
- GPT‑5.5 Instant saw 52% improvement in harm‑to‑others, 39% in suicide/self‑harm.
- Safety summaries are short, factual, limited‑time, used only for serious safety concerns.
- Training involved mental health experts and maintained ordinary conversation quality.
- Evaluated across 4,000+ cases with safety relevance 4.93/5.
- Future work to extend to other high‑risk areas such as biology or cyber safety.
- No meaningful impact on everyday conversation quality.