arXiv:2505.16128cs.CL2025-05ACL被引 5

发现大模型在解题判断中隐含种族偏见,对不同群体解题正确性评价不公。

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning

  • 通过五种主流模型测试,发现模型对不同种族作者的解题结果存在系统性误判。
  • 非洲裔作者的正确解题被低估,亚裔作者在写作评估中得分最低。
  • 模型会无意识为不同群体分配刻板印象颜色,揭示偏见深植于推理过程。

尽管大语言模型在表面层面避免了人口统计学刻板印象,但在多种社会情境下仍表现出偏见。本文发现,大模型在将解题正确性与人口统计学特征关联时存在严重偏差。在数学、编程、常识和写作任务中,对五种人类价值观对齐的大模型进行实验,揭示出两种类型的真实性偏差:归因偏差(Attribution Bias)——模型更倾向于将正确解答归功于特定群体;评估偏差(Evaluation Bias)——对相同解题内容,根据作者的感知身份给出不同评价。结果显示,模型普遍低估非洲裔作者在数学和编程中的正确解答,同时高估其错误解答;而亚裔作者在写作评估中得分最低。额外研究表明,模型在生成可视化代码时会自动为不同群体分配刻板印象颜色,表明这些偏见已深度嵌入模型的推理机制。研究提示,人口统计偏见不仅存在于表层刻板印象或情境诱导中,更广泛渗透于模型的决策过程,对教育与评估场景的应用构成重大风险。

原文摘要 · Abstract (English)

Despite LLMs' explicit alignment against demographic stereotypes, they have been shown to exhibit biases under various social contexts. In this work, we find that LLMs exhibit concerning biases in how they associate solution veracity with demographics. Through experiments across five human value-aligned LLMs on mathematics, coding, commonsense, and writing problems, we reveal two forms of such veracity biases: Attribution Bias, where models disproportionately attribute correct solutions to certain demographic groups, and Evaluation Bias, where models' assessment of identical solutions varies based on perceived demographic authorship. Our results show pervasive biases: LLMs consistently attribute fewer correct solutions and more incorrect ones to African-American groups in math and coding, while Asian authorships are least preferred in writing evaluation. In additional studies, we show LLMs automatically assign racially stereotypical colors to demographic groups in visualization code, suggesting these biases are deeply embedded in models' reasoning processes. Our findings indicate that demographic bias extends beyond surface-level stereotypes and social context provocations, raising concerns about LLMs' deployment in educational and evaluation settings.

大模型偏见认知偏差公平性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。