arXiv:2601.15488cs.CLcs.AI2026-01ACL

通过多角色视角推理,减少大模型社会偏见。

Multi-Persona Thinking for Bias Mitigation in Large Language Models

  • 推理时引入多个社会身份视角协同分析
  • 在两个基准上均低于现有提示方法的偏见水平
  • 适合需要公平性保障的生成应用

大型语言模型存在社会偏见,可能导致有害刻板印象和不公平结果。我们提出多角色思维(Multi-Persona Thinking, MPT),一种简单的推理阶段框架,通过鼓励从多个视角进行推理来减轻社会偏见。MPT引导模型同时考虑对立的社会身份(如男性与女性)以及中立视角,这些视角通过迭代推理过程交互,识别并修正有偏判断。该设计将角色设定的潜在弱点转化为偏见缓解机制。我们在两个常用偏见基准上,对开源与闭源模型进行了评估。结果表明,MPT在保持核心推理能力的同时,偏见水平低于现有提示方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit social biases, which can lead to harmful stereotypes and unfair outcomes. We propose \textbf{Multi-Persona Thinking (MPT)}, a simple inference-time framework that reduces social bias by encouraging reasoning from multiple perspectives. MPT guides the model to consider contrasting social identities, such as male and female, together with a neutral viewpoint. These viewpoints then interact through an iterative reasoning process to identify and correct biased judgments. This design transforms the potential weakness of persona assignment into a mechanism to mitigate bias. We evaluate MPT on two widely used bias benchmarks with both open-source and closed-source models. Our results show that MPT achieves a lower bias than the existing prompting-based methods while maintaining the core reasoning ability.

偏见缓解多角色推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。