arXiv:2503.06011cs.CLcs.AI2025-03被引 6

通过明确意图提升大模型自修正能力,更有效减少社会偏见。

Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models

  • 在指令、回应和反馈中显式表达去偏见意图
  • 多维度反馈结合思维链使偏见减少更稳定
  • 适合关注模型公平性与可解释性的研究者

基于反馈的自修正能提升大语言模型(LLMs)输出质量。从认知心理学看,自修正类似慢速、有意识的系统2思维,可能降低模型的社会偏见。由于LLMs对上下文模糊性和不一致性敏感,因此在自修正过程中显式传达意图至关重要。本研究证明,在自修正各环节——指令、回应与反馈中明确意图,是有效减缓偏见的关键。我们分解自修正为三部分:指令、回应与反馈,并分别进行意图澄清。在指令阶段,引入显式去偏提示;在回应阶段,采用思维链(CoT)揭示推理过程;在反馈阶段,定义去偏评估维度,通过多维度批评与评分提供清晰反馈。实验表明,基于多维度反馈的去偏提示所生成的自修正思维链回应,比基线方法更稳健、一致地减少偏见响应。同时发现,不同偏见水平的模型或分离生成回应与反馈的模型,在去偏效果上存在差异。

原文摘要 · Abstract (English)

Self-Correction based on feedback improves the output quality of Large Language Models (LLMs). Moreover, as Self-Correction functions like the slow and conscious System-2 thinking from cognitive psychology's perspective, it can potentially reduce LLMs' social biases. LLMs are sensitive to contextual ambiguities and inconsistencies; therefore, explicitly communicating their intentions during interactions when applying Self-Correction for debiasing is crucial. In this study, we demonstrate that clarifying intentions is essential for effectively reducing biases in LLMs through Self-Correction. We divide the components needed for Self-Correction into three parts: instruction, response, and feedback, and clarify intentions at each component. We incorporate an explicit debiasing prompt to convey the intention of bias mitigation from the instruction for response generation. In the response, we use Chain-of-Thought (CoT) to clarify the reasoning process. In the feedback, we define evaluation aspects necessary for debiasing and propose clear feedback through multi-aspect critiques and scoring. Through experiments, we demonstrate that self-correcting CoT responses obtained from a debiasing prompt based on multi-aspect feedback can reduce biased responses more robustly and consistently than the baselines. We also find the variation in debiasing efficacy when using models with different bias levels or separating models for response and feedback generation.

去偏见自修正思维链大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。