arXiv:2508.06671cs.CLcs.AI2025-08中稿 · COLM被引 3

研究大模型思考过程是否含偏见,发现其思维与输出偏差关联弱。

Do Biased Models Have Biased Thoughts?

  • 用思维链提示分析模型推理步骤中的偏见
  • 5个主流模型显示思维与输出偏见相关性低于0.6
  • 适合关注AI公平性与可信生成的研究者

语言模型性能卓越,但其在性别、种族、社会经济地位、外貌和性取向等方面的偏见使部署面临挑战。本文研究思维链提示(chain-of-thought prompting)对公平性的影响,核心问题为:‘有偏见的模型是否有偏见的思考?’我们在5个主流大语言模型上使用公平性指标,量化11种不同偏见在模型思考与输出中的表现。结果表明,思维步骤中的偏见与最终输出偏见的相关性较低(多数情况相关性小于0.6,p值小于0.001),说明这些模型即便决策有偏,其推理过程未必伴随偏见,与人类行为模式不同。

原文摘要 · Abstract (English)

The impressive performance of language models is undeniable. However, the presence of biases based on gender, race, socio-economic status, physical appearance, and sexual orientation makes the deployment of language models challenging. This paper studies the effect of chain-of-thought prompting, a recent approach that studies the steps followed by the model before it responds, on fairness. More specifically, we ask the following question: $\textit{Do biased models have biased thoughts}$? To answer our question, we conduct experiments on $5$ popular large language models using fairness metrics to quantify $11$ different biases in the model's thoughts and output. Our results show that the bias in the thinking steps is not highly correlated with the output bias (less than $0.6$ correlation with a $p$-value smaller than $0.001$ in most cases). In other words, unlike human beings, the tested models with biased decisions do not always possess biased thoughts.

语言模型公平性思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。