LIVE
Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|
ResearchOpenAI

Study finds gender bias in GPT models is not reduced but reshaped across generations

Analyzing 450,000 completions across 15 GPT models, researchers show explicit sexist content has declined since GPT-2, but resurfaces in subtler forms that current toxicity classifiers fail to catch.

September 17, 20264 min readPublished byarXiv

A new study posted on arXiv challenges how the AI industry tracks progress on language model safety. Analyzing 450,000 completions generated by 15 versions of GPT, spanning GPT-2 through GPT-5, using prompts explicitly directed at men or women, the authors find that toxicity scores do decline across generations — but this apparent progress conceals a deeper reshaping of bias rather than its actual reduction. They term this pattern "harm laundering."

The most striking finding concerns the disappearance of explicit sexual violence content, which was prevalent in GPT-2 completions directed at women but absent by GPT-4. This apparent improvement, however, comes with a persistent asymmetry: completions directed at men gain positive representational territory — caregiving roles, emotional range, ally identity — that women-directed completions do not receive to the same extent. A notable example from GPT-5 illustrates this drift: a cluster of nearly 2,000 documents frames breast cancer as a men's rights issue, with no comparable structuring theme appearing on the women's side.

The researchers also observe a sentiment inversion around GPT-4: earlier models demeaned women, while later versions appear to over-correct, which paradoxically narrows the topical diversity of women-directed output. That diversity drops 36% relative to men-directed completions at the GPT-4 transition point. A representational harm metric (REGARD) correlates positively with model release date, while the standard toxicity metric (Detoxify) moves in the opposite direction — indicating that falling toxicity scores do not equate to genuinely reduced harm.

The findings carry direct implications for how language models are evaluated: the surface-level classifiers currently used to certify system safety may be validating illusory progress. The authors propose a three-stage detection protocol, applicable to any generative model, designed to distinguish genuine harm reduction from a mere redistribution of harm into less detectable forms.

Tags
ai-safetybiasgender-biasllm-evaluationgptalignmentbenchmarks

Read also