探究思维链提示如何影响大模型性别偏见,发现效果只是表面的。
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs

- 结合基准测试与可解释性分析,研究思维链对偏见的影响机制。
- 思维链未一致降低偏见差距,部分注意力头虽平衡但偏见仍存于隐藏表征。
- 改进源于数据记忆而非真正理解,适合关注模型内部机制的研究者。
大型语言模型在社会敏感场景中日益广泛应用,尽管已有大量文献指出其存在性别偏见。思维链(Chain-of-Thought, CoT)提示被提出作为缓解偏见的方法。然而,现有评估主要关注模型基准性能的变化,难以揭示表观偏见减少是否反映模型内部机制的真实改变。本文结合基准评估、机制可解释性技术及推理链失败分析,研究CoT提示对LLM性别偏见的影响。结果确认了基准测试中普遍存在的刻板偏见,表明CoT提示并未一致减少偏见差距。机制分析显示,尽管某些注意力头簇中的偏见行为得到平衡,但性别偏见仍嵌入在隐藏表征中,说明仅为表面缓解。进一步检查推理链发现,这些改进源于对训练数据的记忆与熟悉度,而非对偏见的真正理解。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in socially sensitive settings despite substantial documentation that they encode gender biases. Chain-of-Thought (CoT) prompting has been proposed as a bias-mitigation approach. However, existing evaluations primarily focus on changes in LLM benchmark performance, providing limited insight into whether apparent bias reductions reflect meaningful changes in a model's internal mechanisms. In this work, we investigate how CoT prompting affects gender bias in LLMs, combining benchmark-based evaluation with mechanistic interpretability techniques and reasoning chain failure analysis. Our results confirm a stereotypical bias present in LLM outputs across benchmarks, showing that CoT prompting does not consistently reduce the bias gap. Mechanistic analyses reveal that although CoT balances biased behavior in certain attention head clusters, gender bias remains embedded in hidden representations, indicating only superficial mitigation. Inspection of reasoning chains further suggests that these improvements stem from memorization and familiarity with the dataset rather than genuine understanding of bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。