通过分析模型内部表示,发现并修正大模型在招聘入学中的种族偏见。
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
- 在模型激活中寻找种族子空间,通过干预实现去偏。
- Gemma模型对白人录取率比黑人高26%,经干预后偏见减少37%-57%。
- 该方法对提示格式敏感,通用种族表示仍难实现。
为提升大语言模型在高风险决策中的公平性,本文构建了基于姓名推断种族的模拟申请场景,用于测试招生与招聘任务中的种族偏见。实验表明,Gemma 2B Instruct 和 LLaMA 3.2 3B Instruct 均存在显著偏见:前者对白人申请者录取率比黑人高出26%,后者对亚裔的雇佣偏好比白人高出60%。多种提示工程策略均无法有效缓解偏见。相反,通过分布式对齐搜索,可在模型激活中识别“种族子空间”,并对其中表示进行平均化干预,使Gemma的偏见降低37%-57%。进一步分析显示,该种族子空间的泛化能力有限,提示格式变化会影响种族表征。结果表明,机制性干预是提升模型公平性的可行路径,但通用种族表示仍难以实现。
原文摘要 · Abstract (English)
Understanding and mitigating biases is critical for the adoption of large language models (LLMs) in high-stakes decision-making. We introduce Admissions and Hiring, decision tasks with hypothetical applicant profiles where a person's race can be inferred from their name, as simplified test beds for racial bias. We show that Gemma 2B Instruct and LLaMA 3.2 3B Instruct exhibit strong biases. Gemma grants admission to 26% more White than Black applicants, and LLaMA hires 60% more Asian than White applicants. We demonstrate that these biases are resistant to prompt engineering: multiple prompting strategies all fail to promote fairness. In contrast, using distributed alignment search, we can identify "race subspaces" within model activations and intervene on them to debias model decisions. Averaging the representation across all races within the subspaces reduces Gemma's bias by 37-57%. Finally, we examine the generalizability of Gemma's race subspaces, and find limited evidence for generalization, where changing the prompt format can affect the race representation. Our work suggests mechanistic approaches may provide a promising venue for improving the fairness of LLMs, but a universal race representation remains elusive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。