arXiv:2602.14529cs.AI2026-02被引 1

区分大模型幻觉与欺骗性错误的内在机制,揭示两者本质不同。

Disentangling Deception and Hallucination Failures in LLMs

  • 从知识存在与行为表达两方面拆解模型失败机理
  • 在受控环境下验证四种行为模式,发现幻觉与欺骗机制迥异
  • 适用于研究模型可信性、安全性和可解释性的研究人员

大型语言模型(LLMs)的错误常从行为层面分析,通常将事实问答中的错误归因于知识缺失。本文聚焦实体类事实问题,提出该视角可能混淆了不同故障机制,主张采用内部机制导向的分析框架,将知识存在性与行为表达分离。在此框架下,幻觉与欺骗属于两种质不同的失败模式,虽输出表现相似,但内在机制不同。为此,我们构建了一个以实体为中心的事实问答受控环境,保持知识完整的同时选择性改变行为表达,实现对四种行为情形的系统分析。通过表示可分性、稀疏可解释性及推理时激活引导等方法,深入剖析这些故障模式。

原文摘要 · Abstract (English)

Failures in large language models (LLMs) are often analyzed from a behavioral perspective, where incorrect outputs in factual question answering are commonly associated with missing knowledge. In this work, focusing on entity-based factual queries, we suggest that such a view may conflate different failure mechanisms, and propose an internal, mechanism-oriented perspective that separates Knowledge Existence from Behavior Expression. Under this formulation, hallucination and deception correspond to two qualitatively different failure modes that may appear similar at the output level but differ in their underlying mechanisms. To study this distinction, we construct a controlled environment for entity-centric factual questions in which knowledge is preserved while behavioral expression is selectively altered, enabling systematic analysis of four behavioral cases. We analyze these failure modes through representation separability, sparse interpretability, and inference-time activation steering.

大模型故障幻觉检测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。