通过隐式辩论分析大模型推理过程,揭示幻觉成因。
Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
- 在单次推理中捕捉模型内部的隐性支持与反驳信号
- 与原模型预测高度一致,且能有效检测幻觉
- 适合研究模型推理机制和幻觉风险的学者
理解大语言模型(LLMs)的内部思考过程及其幻觉成因仍是关键挑战。为此,我们提出隐式辩论(latent debate)框架,通过捕捉单次推理中模型内部隐含的支持与反驳信号,来解释模型预测。不同于依赖多答案或多模型显式辩论的现有方法,该框架不依赖外部代理,而是基于模型自身生成的隐性论证结构。我们首先构建了模型与任务无关的概念框架,并符号化实现,用于近似LLM在真假判断任务中的思考过程。实证研究表明,隐式辩论是原模型的高度一致结构化替代模型,且能作为幻觉检测的强基线。进一步分析发现,幻觉与辩论模式存在显著相关性:中间层隐式辩论程度越高,幻觉风险越大。这些结果表明,隐式辩论可成为理解LLM内部机制的有效框架,尤其适用于推理过程中出现内部(不)一致性的场景。
原文摘要 · Abstract (English)
Understanding the internal thinking process of Large Language Models (LLMs) and the cause of hallucinations remains a key challenge. To this end, we introduce latent debate, a novel framework for interpreting model predictions through the lens of implicit internal arguments. Unlike the current work of self-consistency and multi-agent debate, which relies on explicit debates among multiple answers or multiple models, latent debate captures the hidden supporting and attacking signals that arise within a single model during a single inference. We first present a model- and task-agnostic conceptual framework, and then instantiate it symbolically to approximate the thinking process of LLMs on True/False prediction tasks. Empirical studies demonstrate that latent debate is a faithful structured surrogate model that has highly consistent predictions with the original LLM. Beyond interpretability, we demonstrate that latent debate provides a strong baseline for hallucination detection. Further analysis reveals strong correlations between hallucinations and debate patterns, such as a high degree of latent debates in the middle layers is linked to a higher risk of hallucinations. These findings position latent debate as a potential framework for understanding internal mechanisms of LLMs, especially for scenarios where internal (dis)agreements appear during the inference steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。