Preprint finds detectable occupational bias in language models despite passing behavioral tests
Researchers introduce a causal framework to measure representational bias in LMs, showing demographic attributes influence internal competence representations even when outputs appear unbiased.
1 source · cross-referenced
- A new arXiv preprint introduces a causal framework to separate representational bias from behavioral bias in language models.
- The study finds demographic attributes like gender, race, and socioeconomic status influence models' internal representations of user competence.
- Biases detectable in internal representations can still affect downstream behavior, even when standard behavioral tests show no disparity.
- The framework was tested on multiple open-weight models using question-answering and hiring tasks.
A new arXiv preprint proposes a causal framework to measure representational bias in language models (LMs), arguing that behavioral tests alone may not capture underlying associations that could still influence outputs. The authors, Keren Fuentes and Aaron Mueller, note that LMs often pass behavioral bias evaluations, but it remains unclear whether this indicates a true absence of bias or merely learned suppression of biased expressions.
The framework decomposes occupational bias into two stages: a model's internal representation of a user's competence and its observable outputs. To operationalize this, the researchers derive "steering vectors" for representations of user expertise and verify that these vectors causally mediate model behavior in both a question-answering task and a hiring task.
Applying the framework to several open-weight models, the study finds that demographic attributes such as gender, race, and socioeconomic status influence a model's internal representation of user expertise. Crucially, these representational biases were detectable even in cases where behavioral metrics showed no disparity between demographics.
The authors demonstrate that these internal biases can influence downstream behavior under intervention, suggesting potential failure modes that behavioral metrics alone may fail to detect. This indicates that even models passing standard bias evaluations may still harbor latent representational biases with real-world consequences.
- Aug 24, 2026 · arXiv cs.CL
Paper quantifies ‘clinical lost-in-the-middle’ effect in EHR processing and proposes query-conditioned context selection
Trust79 - Aug 24, 2026 · Apple — Machine Learning Research
Apple proposes Internalized Visual Thinking to cut video reasoning latency by more than fivefold
Trust79 - Aug 24, 2026 · arXiv cs.AI
Paper proposes Spec-Driven Agentic Development as a new paradigm for AI-native software delivery
Trust79