arXiv:2506.10922cs.LGcs.AI2025-06被引 14

用可解释性技术在真实场景中有效降低大模型招聘偏见。

Robustly Improving LLM Fairness in Realistic Settings via Interpretability

  • 通过识别模型激活中的敏感属性方向,实时进行概念编辑来消除偏见。
  • 引入真实上下文后,模型对黑人/女性候选人偏见最高达12%,且普遍存在。
  • 方法通用性强,适用于主流商业与开源模型,保持性能同时将偏见降至1%以下。

大型语言模型(LLMs)越来越多地应用于高风险的招聘场景,直接影响个人职业与生计。尽管以往研究认为简单反偏见提示可在控制环境下消除人口统计偏差,但我们在引入真实情境细节(如公司名称、公开招聘页面的文化描述、仅录用前10%候选人等)后发现,此类缓解措施失效。我们提出内部偏见缓解策略:通过识别并中和模型激活中与种族、性别相关的特征方向,在推理时实施仿射概念编辑。在GPT-4o、Claude 4 Sonnet、Gemini 2.5 Flash、Gemma-2 27B、Gemma-3、Mistral-24B等主流模型上测试表明,加入真实上下文后,种族与性别偏差显著上升(面试率差异最高达12%),且普遍倾向黑人优于白人、女性优于男性。模型甚至能从学院背景等细微线索推断身份,导致偏见难以通过链式思考观察到。使用简单合成数据提取的方向即可实现跨模型泛化,干预后偏见通常低于1%,始终不超过2.5%,同时基本维持模型性能。结果表明,招聘应用中应采用更真实的评估方法,并考虑内嵌缓解机制以保障公平。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in high-stakes hiring applications, making decisions that directly impact people's careers and livelihoods. While prior studies suggest simple anti-bias prompts can eliminate demographic biases in controlled evaluations, we find these mitigations fail when realistic contextual details are introduced. We address these failures through internal bias mitigation: by identifying and neutralizing sensitive attribute directions within model activations, we achieve robust bias reduction across all tested scenarios. Across leading commercial (GPT-4o, Claude 4 Sonnet, Gemini 2.5 Flash) and open-source models (Gemma-2 27B, Gemma-3, Mistral-24B), we find that adding realistic context such as company names, culture descriptions from public careers pages, and selective hiring constraints (e.g.,``only accept candidates in the top 10\%") induces significant racial and gender biases (up to 12\% differences in interview rates). When these biases emerge, they consistently favor Black over White candidates and female over male candidates across all tested models and scenarios. Moreover, models can infer demographics and become biased from subtle cues like college affiliations, with these biases remaining invisible even when inspecting the model's chain-of-thought reasoning. To address these limitations, our internal bias mitigation identifies race and gender-correlated directions and applies affine concept editing at inference time. Despite using directions from a simple synthetic dataset, the intervention generalizes robustly, consistently reducing bias to very low levels (typically under 1\%, always below 2.5\%) while largely maintaining model performance. Our findings suggest that practitioners deploying LLMs for hiring should adopt more realistic evaluation methodologies and consider internal mitigation strategies for equitable outcomes.

大模型公平性招聘应用可解释性偏见缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。