arXiv:2604.11087cs.LG2026-04ACL被引 1

用反事实干预识别大模型幻觉,提升可信度

CausalGaze: Unveiling Hallucinations via Counterfactual Graph Intervention in Large Language Models

  • 构建动态因果图,通过反事实干预分离真实推理路径
  • 在TruthfulQA上比现有方法提升3.3%的AUROC
  • 适合关注模型可解释性与高风险场景应用的研究者

尽管大语言模型取得了突破性进展,但幻觉仍是其在高风险领域部署的关键瓶颈。现有基于分类的方法主要依赖内部状态的静态、被动信号,常捕捉到噪声和虚假关联,忽视了潜在的因果机制。为此,我们提出从被动观测转向主动干预的新范式,引入基于结构因果模型(SCMs)的幻觉检测框架CausalGaze。该方法将大模型内部状态建模为动态因果图,并利用反事实干预来剥离因果推理路径中的偶然噪声,从而增强模型可解释性。在四个数据集和三种主流大模型上的大量实验表明,CausalGaze效果显著,尤其在TruthfulQA数据集上相比最先进基线提升3.3%的AUROC。

原文摘要 · Abstract (English)

Despite the groundbreaking advancements made by large language models (LLMs), hallucination remains a critical bottleneck for their deployment in high-stakes domains. Existing classification-based methods mainly rely on static and passive signals from internal states, which often captures the noise and spurious correlations, while overlooking the underlying causal mechanisms. To address this limitation, we shift the paradigm from passive observation to active intervention by introducing CausalGaze, a novel hallucination detection framework based on structural causal models (SCMs). CausalGaze models LLMs' internal states as dynamic causal graphs and employs counterfactual interventions to disentangle causal reasoning paths from incidental noise, thereby enhancing model interpretability. Extensive experiments across four datasets and three widely used LLMs demonstrate the effectiveness of CausalGaze, especially achieving 3.3% improvement in AUROC on the TruthfulQA dataset compared to state-of-the-art baselines.

幻觉检测因果模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。