语言模型内部仍存职业能力偏见,即使表面表现无差异。
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

- 用因果框架拆解偏见:区分内部表征与输出行为
- 发现性别、种族等影响模型对用户专业度的内部判断
- 揭示行为指标无法检测的隐藏偏见,适合安全评估者关注
语言模型在行为测试中常表现无偏,但其内部是否存在偏见仍不明确。本研究提出一种因果分析框架,将职业偏见分解为两个维度:模型对用户专业能力的内部表征,以及可观察的输出行为。通过构建用户专业度的引导向量,验证其在问答和招聘任务中对模型行为的因果影响。在多个开源模型上测试发现,性别、种族、社会经济地位等人口属性仍会影响模型对用户专业性的内部表征,即便行为指标未显示群体差异。干预实验表明这些内部表征能影响下游决策,提示仅依赖行为评估可能遗漏关键风险。
原文摘要 · Abstract (English)
Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。