arXiv:2601.01552cs.CL2026-01Conference of the …被引 4

用拓扑分析注意力动态,识别大模型幻觉生成

HalluZig: Hallucination Detection using Zigzag Persistence

  • 通过注意力矩阵的拓扑演变捕捉内部推理缺陷
  • 在多个基准上超越强基线,准确率显著提升
  • 仅需部分网络层结构即可通用检测,适用性强

大型语言模型(LLM)因容易产生幻觉,其事实可靠性仍是高风险领域应用的关键障碍。现有检测方法多依赖输出表面信号,忽视模型内部推理过程中的错误。本文提出一种新范式:通过分析模型各层注意力随时间演化的动态拓扑,构建注意力矩阵的锯齿图过滤,利用拓扑数据分析中的锯齿持久性提取拓扑签名。核心假设为:真实生成与幻觉生成具有不同的拓扑特征。我们在多个基准上验证了该框架 HalluZig,结果表明其性能优于多个强基线。此外,分析显示这些拓扑签名在不同模型间具备泛化能力,仅使用部分网络深度的结构特征即可实现幻觉检测。

原文摘要 · Abstract (English)

The factual reliability of Large Language Models (LLMs) remains a critical barrier to their adoption in high-stakes domains due to their propensity to hallucinate. Current detection methods often rely on surface-level signals from the model's output, overlooking the failures that occur within the model's internal reasoning process. In this paper, we introduce a new paradigm for hallucination detection by analyzing the dynamic topology of the evolution of model's layer-wise attention. We model the sequence of attention matrices as a zigzag graph filtration and use zigzag persistence, a tool from Topological Data Analysis, to extract a topological signature. Our core hypothesis is that factual and hallucinated generations exhibit distinct topological signatures. We validate our framework, HalluZig, on multiple benchmarks, demonstrating that it outperforms strong baselines. Furthermore, our analysis reveals that these topological signatures are generalizable across different models and hallucination detection is possible only using structural signatures from partial network depth.

幻觉检测拓扑分析大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。