arXiv:2508.07753cs.CL2025-08中稿 · CIKM 2025

发现社会偏见是导致大模型幻觉的关键原因,且影响方向因偏见类型而异。

Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models

  • 用因果模型和干预数据集分离偏见与幻觉的因果关系
  • 实验表明偏见显著引发不忠实幻觉,不同偏见影响方向各异
  • 适合关注模型公平性与可信度的研究者阅读

大型语言模型(LLMs)在多项任务中表现卓越,但仍存在输出与输入不符的忠实性幻觉问题。本文首次探究社会偏见是否构成此类幻觉的因果因素。核心挑战在于上下文中的混杂变量难以控制,影响因果推断。为此,我们采用结构因果模型(SCM)建立并验证因果关系,并设计偏见干预方法以控制混杂因子。同时,构建了包含多种社会偏见的偏见干预数据集(BID),实现对因果效应的精确测量。在主流大模型上的实验表明,社会偏见是忠实性幻觉的重要成因,且每种偏见状态的影响方向不同。进一步分析显示,偏见主要诱发不公平幻觉,在不同模型间存在细微但显著的因果效应。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable success in various tasks, yet they remain vulnerable to faithfulness hallucinations, where the output does not align with the input. In this study, we investigate whether social bias contributes to these hallucinations, a causal relationship that has not been explored. A key challenge is controlling confounders within the context, which complicates the isolation of causality between bias states and hallucinations. To address this, we utilize the Structural Causal Model (SCM) to establish and validate the causality and design bias interventions to control confounders. In addition, we develop the Bias Intervention Dataset (BID), which includes various social biases, enabling precise measurement of causal effects. Experiments on mainstream LLMs reveal that biases are significant causes of faithfulness hallucinations, and the effect of each bias state differs in direction. We further analyze the scope of these causal effects across various models, specifically focusing on unfairness hallucinations, which are primarily targeted by social bias, revealing the subtle yet significant causal effect of bias on hallucination generation.

大模型偏见幻觉因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。