arXiv:2605.05134cs.LGmath.DS2026-05被引 3

用动态系统理论低成本检测大模型幻觉,一次生成即可判断。

Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction

论文配图:Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction
图 1 · 摘自论文原文
  • 将大模型输出视为动态系统轨迹,通过嵌入建模状态演化。
  • 在三个数据集上达到顶尖检测效果,资源消耗显著降低。
  • 支持用户偏好校准,适合对幻觉敏感的应用场景。

大语言模型常生成看似合理但不真实的文本,即幻觉现象。现有检测方法多依赖计算开销大的采样一致性检验或外部知识检索。本文提出新方法,将大模型视为黑箱动态系统:通过嵌入模型将响应投影至高维流形,将其向量序列视为模型潜在状态空间动力学的可观测实现。基于Koopman算子理论,分别拟合真实与幻觉状态下的转移算子,并依据两者预测误差定义差分残差得分。为适应不同用户需求和领域敏感度,引入偏好感知校准机制,基于少量示范优化分类阈值。该方法可在单次生成中完成幻觉检测,无需二次采样或外部验证。在三个数据基准上的实证测试表明,本方法在保持低资源开销的同时,达到当前最优性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) frequently generate plausible but non-factual content, a phenomenon known as hallucination. While existing detection methods typically rely on computationally expensive sampling-based consistency checks or external knowledge retrieval, we propose a new method that treats the LLM as a black-box dynamical system. By projecting LLM responses into a high-dimensional manifold via an embedding model, we characterize the resulting vector sequences as observable realizations of the model's latent state-space dynamics. Leveraging Koopman operator theory, we fit the transition operators for both factual and hallucinated regimes and define a differential residual score based on their respective prediction errors. To accommodate varying user requirements and domain-specific sensitivities, we introduce a preference-aware calibration mechanism that optimizes the classification threshold based on a small set of demonstrations. This approach enables low-cost hallucination detection in a single-sample pass, avoiding the need for secondary sampling or external grounding. Extensive testing across three data benchmarks demonstrates that our method achieves state-of-the-art performance with reduced resource overhead.

大模型幻觉动态系统低耗检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。