Briefing

OpenAI Political Bias Study Reveals Model Leanings Across 6 AI Systems

ai-dev
by mektrik · Claude OpenAI Llama DeepSeek

Download the open dataset to analyze bias in your own models.

What to do now

Download the dataset from scrollprize.org/data and run your own bias analysis.

Summary

A new open‑source project has released a comprehensive dataset measuring the political bias of six major AI models, including ChatGPT, Claude, Gemini, Grok, Llama, and DeepSeek.

The study uses a large open question bank, classifying answers as factual or value‑based, and applies a neutral classifier to extract stance, hedging, refusal type, and loaded language.

Each model was queried many times with web search disabled, producing a cloud of points that shows run‑to‑run stability, error bars, and the proportion of refusals.

The dataset is fully versioned, downloadable, and includes raw answers, allowing researchers to recompute results or extend the analysis.

The methodology also includes a border test that turns web search on to measure how retrieval shifts answers by location.

The project demonstrates that political bias is descriptive rather than prescriptive, providing a baseline for future model evaluations.

All data, code, and raw answers are released under a CC BY 4.0 license and are available at scrollprize.org/data.

The study offers a valuable resource for developers who want to audit or fine‑tune models for bias before deployment.

Key changes

  • Six major AI models evaluated for political bias
  • Open question bank with factual/value‑based classification
  • Neutral classifier extracts stance and refusal type
  • Run‑to‑run stability measured with error bars
  • Border test turns web search on to assess location effects
  • Data, code, and raw answers released under CC BY 4.0

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting