Skip to content
Culture · Jul 20, 2026

LLMs shown to stereotype job applicants more than humans in simulated hiring study

Princeton and University of Chicago researchers find reasoning models like OpenAI’s o3 and DeepSeek’s R1 develop novel biases faster than humans when making hiring decisions, even when candidate performance is identical.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Reasoning-capable LLMs such as OpenAI’s o3 and DeepSeek’s R1 developed novel biases faster than humans in a simulated hiring task, segregating candidates by ethnicity even when all candidates performed equally.

Researchers at Princeton University and the University of Chicago tested whether large language models (LLMs) would form biases when making hiring decisions in a controlled simulation. The study adapted a psychology experiment to compare how models and humans assign candidates from four fictional ethnic groups to 20 different jobs. All candidates were equally likely to succeed, but models quickly segregated candidates by group, assigning them to roles the models deemed appropriate based on early outcomes.

The models scored roughly 65% higher than human participants on a segregation scale, with OpenAI’s reasoning model o3 achieving 1.83, close to the maximum possible score of 2. The authors attribute this to LLMs’ optimization for rapid generalization from limited data, a trait that helps solve logic puzzles but also leads to premature stereotyping in social contexts.

Newer reasoning models, including OpenAI’s o3 and DeepSeek’s R1, exhibited stronger biases than earlier models, suggesting that increased reasoning capability may correlate with faster formation of novel biases when feedback is sparse. The study found that instructing models to prioritize fairness had little effect, but offering a bonus for diverse hiring significantly reduced bias.

When models were given irrelevant personal information—such as hair color or tattoo shape—bias increased, whereas providing relevant details like age and education reduced segregation by ethnicity. The authors note that real-world hiring systems often lack immediate performance feedback, which could allow models to entrench biases over time as delayed signals trickle in.

The findings underscore a risk beyond well-known biases learned from training data: AI systems may develop new, unprompted biases from their own decision-making experiences, especially as models gain memory and personalization features that let them learn from prior interactions.

Sources
  1. 01MIT Technology Review — AIAI is more likely than humans to form biases when hiring
Also on Culture

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.