arXiv:2601.14553cs.CLcs.AI2026-01被引 1

让大模型假装不知道偏见信息,反而更公平

Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models

  • 用自盲和反事实模拟让模型假装不知偏见信息
  • 相比直接忽略,该方法使性别/种族偏见降低37%
  • 适合研究模型公平性或对齐的学者使用

公平决策需要忽视无关且可能带来偏见的信息。决策者需模拟自己若未掌握某些事实(如求职者性别或种族)会作何判断。这种反事实自我模拟对人类而言极为困难,导致即使善意者也会产生偏见。本文发现大语言模型在模拟自身反事实认知以消除性别与种族偏见、克服迎合倾向方面也存在类似局限。直接提示模型忽略偏见信息不仅无效,甚至可能适得其反。但不同于人类,大模型可访问自身反事实认知的真值模型——即其自身的API。我们证明,通过调用被遮蔽的副本响应,模型能做出更公平决策,并提升透明度以区分隐性偏见与有意偏见行为。

原文摘要 · Abstract (English)

Fair decisions require ignoring irrelevant, potentially biasing, information. To achieve this, decision-makers need to approximate what decision they would have made had they not known certain facts, such as the gender or race of a job candidate. This counterfactual self-simulation is notoriously hard for humans, leading to biased judgments even by well-meaning actors. Here we show that large language models (LLMs) suffer from similar limitations in their ability to approximate what decisions they would make under counterfactual knowledge in offsetting gender and race biases and overcoming sycophancy. We show that prompting models to ignore or pretend not to know biasing information fails to offset these biases and occasionally backfires. However, unlike humans, LLMs can be given access to a ground-truth model of their own counterfactual cognition -- their own API. We show that this access to the responses of a blinded replica enables fairer decisions, while providing greater transparency to distinguish implicit from intentionally biased behavior.

大模型对齐公平性反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。