arXiv:2509.10753cs.LGcs.AI2025-09被引 2

用物理场理论分析大模型输出,识别幻觉内容。

HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling

  • 将大模型输出建模为能量与熵的路径场,通过温度变化观察稳定性。
  • 无需微调,在多个数据集上达到当前最佳幻觉检测效果。
  • 适合关注大模型可信度、需无额外训练的工程部署者。

大语言模型(LLMs)虽具备强大的推理和问答能力,但常产生不准确或不可靠的内容,即幻觉,严重限制其在高风险场景的应用。本文提出霍鲁菲尔德(HalluField),一种基于参数化变分原理和热力学的新型幻觉检测方法。受热力学启发,该方法将模型在特定查询与温度设置下的响应视为一系列离散的概率词元路径,每条路径关联对应能量与熵。通过分析温度与似然性变化下能量与熵分布的动态,量化响应的语义稳定性。幻觉通过能量景观中的不稳定或异常行为来识别。该方法计算高效且实用:直接作用于模型输出的logits,无需微调或附加神经网络。方法具有坚实的物理基础,类比热力学第一定律。令人瞩目的是,通过物理视角建模,霍鲁菲尔德在多种模型和数据集上均实现领先性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit impressive reasoning and question-answering capabilities. However, they often produce inaccurate or unreliable content known as hallucinations. This unreliability significantly limits their deployment in high-stakes applications. Thus, there is a growing need for a general-purpose method to detect hallucinations in LLMs. In this work, we introduce HalluField, a novel field-theoretic approach for hallucination detection based on a parametrized variational principle and thermodynamics. Inspired by thermodynamics, HalluField models an LLM's response to a given query and temperature setting as a collection of discrete likelihood token paths, each associated with a corresponding energy and entropy. By analyzing how energy and entropy distributions vary across token paths under changes in temperature and likelihood, HalluField quantifies the semantic stability of a response. Hallucinations are then detected by identifying unstable or erratic behavior in this energy landscape. HalluField is computationally efficient and highly practical: it operates directly on the model's output logits without requiring fine-tuning or auxiliary neural networks. Notably, the method is grounded in a principled physical interpretation, drawing analogies to the first law of thermodynamics. Remarkably, by modeling LLM behavior through this physical lens, HalluField achieves state-of-the-art hallucination detection performance across models and datasets.

幻觉检测大模型热力学可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。