通过结构熵定位并阻断大模型幻觉传播路径,实现精准干预。
LANCET: Neural Intervention via Structural Entropy for Mitigating Faithfulness Hallucinations in LLMs
- 基于梯度对比分析定位易产生幻觉的神经元。
- 通过最小化结构熵追踪幻觉传播路径,实现精准定位。
- 分级干预策略在抑制幻觉的同时保留模型通用能力。
大型语言模型虽革新了信息处理,但其可靠性因幻觉问题严重受损。现有方法多依赖节点级调整或粗粒度抑制,忽视了神经信息的分布式特性,导致干预不精准。我们观察到幻觉传播如同感染,沿特定前向路径扩散。为此提出Lancet框架,利用结构熵与幻觉差异比实现精准神经干预:首先通过梯度驱动的对比分析定位幻觉高风险神经元,再以最小化结构熵映射其传播路径,最后实施分级干预策略,保障模型整体能力。在多个幻觉基准数据集上的综合评估显示,Lancet显著优于当前最优方法,验证了该手术式干预的有效性。
原文摘要 · Abstract (English)
Large Language Models have revolutionized information processing, yet their reliability is severely compromised by faithfulness hallucinations. While current approaches attempt to mitigate this issue through node-level adjustments or coarse suppression, they often overlook the distributed nature of neural information, leading to imprecise interventions. Recognizing that hallucinations propagate through specific forward transmission pathways like an infection, we aim to surgically block this flow using precise structural analysis. To leverage this, we propose Lancet, a novel framework that achieves precise neural intervention by leveraging structural entropy and hallucination difference ratios. Lancet first locates hallucination-prone neurons via gradient-driven contrastive analysis, then maps their propagation pathways by minimizing structural entropy, and finally implements a hierarchical intervention strategy that preserves general model capabilities. Comprehensive evaluations across hallucination benchmark datasets demonstrate that Lancet significantly outperforms state-of-the-art methods, validating the effectiveness of our surgical approach to neural intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。