arXiv:2506.11088cs.CLcs.AI2025-06NeurIPS

统一修复大模型的虚构与失真幻觉,效果更优且不互相干扰。

One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs

  • 通过分析神经激活空间,发现两类幻觉共享部分表征子空间。
  • 提出SPACE框架,联合编辑共享子空间,提升事实性和忠实性。
  • 适合关注大模型可靠性、部署安全性的研究人员和工程师。

大语言模型在自然语言处理中展现出前所未有的能力,但其实际应用仍受持续存在的事实性与忠实性幻觉阻碍。现有方法分别应对这两类幻觉,却常引发性能权衡——针对一类的干预反而加剧另一类。通过对大模型激活空间动态的实证与理论分析,我们发现这两类幻觉在神经表征中存在重叠子空间,为同时缓解提供了可能。为此,我们提出SPACE统一框架,通过双任务特征建模建立共享子空间的几何基础,并结合谱聚类与注意力头显著性评分的混合探针策略识别并编辑这些子空间。在多个基准数据集上的实验结果表明,该方法显著优于现有方案。

原文摘要 · Abstract (English)

LLMs have demonstrated unprecedented capabilities in natural language processing, yet their practical deployment remains hindered by persistent factuality and faithfulness hallucinations. While existing methods address these hallucination types independently, they inadvertently induce performance trade-offs, as interventions targeting one type often exacerbate the other. Through empirical and theoretical analysis of activation space dynamics in LLMs, we reveal that these hallucination categories share overlapping subspaces within neural representations, presenting an opportunity for concurrent mitigation. To harness this insight, we propose SPACE, a unified framework that jointly enhances factuality and faithfulness by editing shared activation subspaces. SPACE establishes a geometric foundation for shared subspace existence through dual-task feature modeling, then identifies and edits these subspaces via a hybrid probe strategy combining spectral clustering and attention head saliency scoring. Experimental results across multiple benchmark datasets demonstrate the superiority of our approach.

大模型幻觉一致性优化神经表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。