剖析大模型生成社会偏见的因果推理机制,发现其常误把相关当因果。
BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models
- 构建三类因果推理分类框架,系统测试模型对敏感议题的推理逻辑
- 4个主流模型在1788个问题中普遍表现出偏见性因果推理
- 揭示模型会先误判相关性为因果,再输出偏见结论,可为去偏提供方向
尽管大语言模型在社会中作用日益重要,研究仍表明它们持续生成针对敏感群体的社会偏见内容。现有评测能识别偏见,但难以揭示其背后的推理过程。本文通过评估大模型回答社会偏见问题时的因果推理,提出一个形式化分类框架,将因果推理分为三类:错误、偏见和情境化。我们合成1788个问题,覆盖八类敏感属性,每组问题专门探测特定推理类型。所有问题均经人工验证,并要求模型生成支撑答案的因果图。评估四个前沿大模型发现,所有模型在多数引发偏见推理的问题上均表现偏见。此外,我们发现模型普遍存在“错误-偏见”推理:先误将相关性当作因果推断敏感群体身份,再进行偏见推理。通过分析模型产生无偏推理的案例,我们识别出三种去偏策略——明确拒绝回答、回避敏感属性、添加情境限制,为未来去偏研究提供思路。
原文摘要 · Abstract (English)
While large language models (LLMs) play increasingly significant roles in society, research shows they continue to generate content that reflects social bias against sensitive groups. Existing benchmarks effectively identify these biases, but a critical gap remains in understanding the underlying reasoning processes that produce them. This paper addresses this gap by evaluating the causal reasoning of LLMs when answering socially biased questions. We propose a formal schema that categorizes causal reasoning into three types (mistaken, biased, and contextually-grounded). We then synthesize 1788 questions covering eight sensitive attributes, with each set of questions designed to probe a specific type of causal reasoning. All questions are then manually validated, and each of them prompts the LLM to generate a causal graph behind its answer. We evaluate four state-of-the-art LLMs and find that all models exhibit biased causal reasoning on most questions eliciting it. Moreover, we discover that LLMs are also prone to "mistaken-biased" reasoning, where they first confuse correlation with causality to infer sensitive group membership and subsequently apply biased causal reasoning. By examining the cases where LLMs produce unbiased causal reasoning, we also identify three strategies LLMs employ to avoid bias (i.e., explicitly refusing to answer, avoiding sensitive attributes, and adding contextual restrictions), which provide insights for future debiasing efforts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。