发现推理模型会重复刻板印象并捏造信息,提出轻量提示法有效降偏。
Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
- 分析推理过程中的两种偏见成因:重复刻板印象、捏造无关细节。
- 在BBQ、StereoSet和BOLD测试中,偏见降低且准确率不降反升。
- 适合关注大模型社会偏见与可解释性研究的读者。
尽管基于推理的大语言模型通过内部结构化思维过程在复杂任务上表现优异,但其思维过程可能累积社会刻板印象,导致偏见输出。然而,这类模型在社会偏见场景下的内在行为仍缺乏深入研究。本文系统探究了这一现象背后的思维机制,揭示出两种导致偏见聚合的失败模式:1)刻板印象重复,即模型主要依赖社会刻板印象作为论证依据;2)无关信息注入,即模型虚构或引入新细节以支撑偏见叙事。基于此,我们提出一种轻量级提示式缓解方法,引导模型自我审查初始推理是否符合上述两类错误。在问答(BBQ、StereoSet)与开放式生成(BOLD)基准上的实验表明,该方法能有效降低偏见,同时保持甚至提升准确性。
原文摘要 · Abstract (English)
While reasoning-based large language models excel at complex tasks through an internal, structured thinking process, a concerning phenomenon has emerged that such a thinking process can aggregate social stereotypes, leading to biased outcomes. However, the underlying behaviours of these language models in social bias scenarios remain underexplored. In this work, we systematically investigate mechanisms within the thinking process behind this phenomenon and uncover two failure patterns that drive social bias aggregation: 1) stereotype repetition, where the model relies on social stereotypes as its primary justification, and 2) irrelevant information injection, where it fabricates or introduces new details to support a biased narrative. Building on these insights, we introduce a lightweight prompt-based mitigation approach that queries the model to review its own initial reasoning against these specific failure patterns. Experiments on question answering (BBQ and StereoSet) and open-ended (BOLD) benchmarks show that our approach effectively reduces bias while maintaining or improving accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。