研究大模型在归因判断中的偏见,揭示其对不同群体的不公平倾向。
Talent or Luck? Evaluating Attribution Bias in Large Language Models
- 基于认知心理学构建归因偏见评估框架
- 发现模型对不同人群的成败归因存在系统性偏差
- 适合关注AI公平性与社会影响的研究者
当学生考试失败时,我们更倾向于归因于个人努力不足还是试题难度过高?归因是指对事件结果的原因分配,它影响认知、强化刻板印象并决定决策。社会心理学中的归因理论解释了人类如何通过隐性认知将原因归为内部(如努力、能力)或外部(如任务难度、运气)因素。大语言模型对事件结果的归因若涉及人口属性,将带来重要的公平性问题。现有研究多聚焦表面关联或孤立刻板印象。本文提出一个基于认知心理学的偏见评估框架,以识别模型推理差异如何导致对不同人口群体的偏见传递。
原文摘要 · Abstract (English)
When a student fails an exam, do we tend to blame their effort or the test's difficulty? Attribution, defined as how reasons are assigned to event outcomes, shapes perceptions, reinforces stereotypes, and influences decisions. Attribution Theory in social psychology explains how humans assign responsibility for events using implicit cognition, attributing causes to internal (e.g., effort, ability) or external (e.g., task difficulty, luck) factors. LLMs' attribution of event outcomes based on demographics carries important fairness implications. Most works exploring social biases in LLMs focus on surface-level associations or isolated stereotypes. This work proposes a cognitively grounded bias evaluation framework to identify how models' reasoning disparities channelize biases toward demographic groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。