用动力系统理论检测大模型幻觉,找知识边界上的不稳定性。
Lyapunov Probes for Hallucination Detection in Large Foundation Models
- 将大模型视为动态系统,用稳定平衡点表示真实知识
- 在知识过渡区边界处检测到幻觉,准确率显著提升
- 适合关注模型可信度与安全性的研究者
我们通过动力系统稳定性理论框架解决大语言模型(LLMs)和多模态大语言模型(MLLMs)的幻觉检测问题。不将幻觉视为简单分类任务,而是将(M)LLMs视为动态系统,其中事实知识由表示空间中的稳定平衡点体现。核心洞察是:幻觉往往出现在分隔稳定区与不稳定区的知识过渡区域边界。为此,我们提出Lyapunov Probes——一种轻量级网络,通过基于导数的稳定性约束训练,强制在输入扰动下置信度单调下降。通过系统的扰动分析与两阶段训练,该方法能可靠区分稳定的真实知识区与易产生幻觉的不稳定区。在多种数据集和模型上的实验表明,其性能持续优于现有基线。
原文摘要 · Abstract (English)
We address hallucination detection in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) by framing the problem through the lens of dynamical systems stability theory. Rather than treating hallucination as a straightforward classification task, we conceptualize (M)LLMs as dynamical systems, where factual knowledge is represented by stable equilibrium points within the representation space. Our main insight is that hallucinations tend to arise at the boundaries of knowledge-transition regions separating stable and unstable zones. To capture this phenomenon, we propose Lyapunov Probes: lightweight networks trained with derivative-based stability constraints that enforce a monotonic decay in confidence under input perturbations. By performing systematic perturbation analysis and applying a two-stage training process, these probes reliably distinguish between stable factual regions and unstable, hallucination-prone regions. Experiments on diverse datasets and models demonstrate consistent improvements over existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。