A new study posted on arXiv challenges how the AI industry tracks progress on language model safety. Analyzing 450,000 completions generated by 15 versions of GPT, spanning GPT-2 through GPT-5, using prompts explicitly directed at men or women, the authors find that toxicity scores do decline across generations — but this apparent progress conceals a deeper reshaping of bias rather than its actual reduction. They term this pattern "harm laundering."
The most striking finding concerns the disappearance of explicit sexual violence content, which was prevalent in GPT-2 completions directed at women but absent by GPT-4. This apparent improvement, however, comes with a persistent asymmetry: completions directed at men gain positive representational territory — caregiving roles, emotional range, ally identity — that women-directed completions do not receive to the same extent. A notable example from GPT-5 illustrates this drift: a cluster of nearly 2,000 documents frames breast cancer as a men's rights issue, with no comparable structuring theme appearing on the women's side.
The researchers also observe a sentiment inversion around GPT-4: earlier models demeaned women, while later versions appear to over-correct, which paradoxically narrows the topical diversity of women-directed output. That diversity drops 36% relative to men-directed completions at the GPT-4 transition point. A representational harm metric (REGARD) correlates positively with model release date, while the standard toxicity metric (Detoxify) moves in the opposite direction — indicating that falling toxicity scores do not equate to genuinely reduced harm.
The findings carry direct implications for how language models are evaluated: the surface-level classifiers currently used to certify system safety may be validating illusory progress. The authors propose a three-stage detection protocol, applicable to any generative model, designed to distinguish genuine harm reduction from a mere redistribution of harm into less detectable forms.