arXiv:2412.10246cs.LG2024-12中稿 · EMNLP被引 9

通过分析模型层间信息流动,检测大模型在模糊提示下的幻觉问题。

Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions

  • 从层间信息传输中捕捉幻觉信号,而非仅看最终输出。
  • 揭示了模型在处理模糊输入时的信息缺失现象。
  • 无需训练或修改结构,可直接集成到现有大模型中。

大型语言模型(LLMs)经常生成看似自信但不准确的回答,这在安全关键领域部署时带来重大风险。本文提出一种新的测试阶段方法,通过系统分析模型各层间的信息流动来检测模型幻觉。重点关注模型处理具有模糊或信息不足上下文的输入时的情况。研究发现,幻觉表现为层间传输中的可用信息不足。与现有方法主要关注最终层输出不同,我们证明追踪跨层信息动态($/mathcal{L}$I)能提供稳健的模型可靠性指标,同时考虑计算过程中的信息增益与损失。$/mathcal{L}$I可轻松集成到预训练的LLM中,无需额外训练或架构修改。

原文摘要 · Abstract (English)

Large language models (LLMs) frequently generate confident yet inaccurate responses, introducing significant risks for deployment in safety-critical domains. We present a novel, test-time approach to detecting model hallucination through systematic analysis of information flow across model layers. We target cases when LLMs process inputs with ambiguous or insufficient context. Our investigation reveals that hallucination manifests as usable information deficiencies in inter-layer transmissions. While existing approaches primarily focus on final-layer output analysis, we demonstrate that tracking cross-layer information dynamics ($\mathcal{L}$I) provides robust indicators of model reliability, accounting for both information gain and loss during computation. $\mathcal{L}$I integrates easily with pretrained LLMs without requiring additional training or architectural modifications.

幻觉检测层间分析大模型可靠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。