arXiv:2506.00088cs.CLcs.AI2025-06ACL被引 8

用微分方程捕捉大模型隐空间动态,提升幻觉检测准确性

HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs

  • 基于神经微分方程建模大模型隐空间演化过程
  • 在五个数据集上比顶尖方法提升超14%的AUC-ROC
  • 适合关注生成可靠性与可信推理的研究者

近年来,大语言模型(LLMs)取得显著进展,但生成虚假或非事实内容的幻觉问题仍是实际部署中的主要挑战。现有基于分类的方法(如SAPLMA)虽高效,但在输出序列早期或中期出现非事实信息时效果下降。为此,我们提出霍尔蒙幻觉检测神经微分方程(HD-NDEs),通过捕捉大模型在隐空间中的完整动态过程来系统评估陈述真实性。该方法利用神经微分方程(Neural DEs)建模大模型隐空间中的动态系统,并将隐空间序列映射至分类空间进行真伪判断。在五个数据集和六种主流大模型上的广泛实验表明,该方法效果显著,尤其在True-False数据集上比现有最优技术提升超过14%的AUC-ROC。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) have made remarkable advancements, yet hallucination, where models produce inaccurate or non-factual statements, remains a significant challenge for real-world deployment. Although current classification-based methods, such as SAPLMA, are highly efficient in mitigating hallucinations, they struggle when non-factual information arises in the early or mid-sequence of outputs, reducing their reliability. To address these issues, we propose Hallucination Detection-Neural Differential Equations (HD-NDEs), a novel method that systematically assesses the truthfulness of statements by capturing the full dynamics of LLMs within their latent space. Our approaches apply neural differential equations (Neural DEs) to model the dynamic system in the latent space of LLMs. Then, the sequence in the latent space is mapped to the classification space for truth assessment. The extensive experiments across five datasets and six widely used LLMs demonstrate the effectiveness of HD-NDEs, especially, achieving over 14% improvement in AUC-ROC on the True-False dataset compared to state-of-the-art techniques.

幻觉检测神经微分方程大模型可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。