arXiv:2603.15557cs.CV2026-03被引 1

通过认知轨迹分析,精准定位视觉语言模型的幻觉生成根源。

Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models

  • 将幻觉视为认知过程的动态异常,构建可解释的低维认知状态空间。
  • 发现几何异常与信息突变本质等价,实现高精度幻觉检测。
  • 可追溯感知、推理、决策三类病理状态,适合可信AI研发者使用。

视觉语言模型(VLMs)常产生看似合理却事实错误的输出,严重阻碍其可信部署。本文提出一种新诊断范式,将幻觉从静态输出错误转化为模型计算认知的动态病理。基于计算理性准则,将生成过程建模为动态认知轨迹,并设计信息论探针将其投影至可解释的低维认知状态空间。核心发现为‘几何-信息对偶性’:轨迹几何异常等价于高信息论意外。幻觉检测转化为几何异常检测。在多种场景下验证——包括严格二元问答(POPE)、综合推理(MME)及开放生成(MS-COCO)——均达当前最优性能。该方法在弱监督下高效运行,且在校准数据严重污染时仍保持鲁棒。能因果归因失败,量化感知不稳(感知熵)、逻辑因果失效(推断冲突)、决策模糊(决策熵)三类病理状态。为构建可透明、可审计、可诊断的AI系统开辟新路径。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) frequently "hallucinate" - generate plausible yet factually incorrect statements - posing a critical barrier to their trustworthy deployment. In this work, we propose a new paradigm for diagnosing hallucinations, recasting them from static output errors into dynamic pathologies of a model's computational cognition. Our framework is grounded in a normative principle of computational rationality, allowing us to model a VLM's generation as a dynamic cognitive trajectory. We design a suite of information-theoretic probes that project this trajectory onto an interpretable, low-dimensional Cognitive State Space. Our central discovery is a governing principle we term the geometric-information duality: a cognitive trajectory's geometric abnormality within this space is fundamentally equivalent to its high information-theoretic surprisal. Hallucination detection is counts as a geometric anomaly detection problem. Evaluated across diverse settings - from rigorous binary QA (POPE) and comprehensive reasoning (MME) to unconstrained open-ended captioning (MS-COCO) - our framework achieves state-of-the-art performance. Crucially, it operates with high efficiency under weak supervision and remains highly robust even when calibration data is heavily contaminated. This approach enables a causal attribution of failures, mapping observable errors to distinct pathological states: perceptual instability (measured by Perceptual Entropy), logical-causal failure (measured by Inferential Conflict), and decisional ambiguity (measured by Decision Entropy). Ultimately, this opens a path toward building AI systems whose reasoning is transparent, auditable, and diagnosable by design.

幻觉检测认知建模VLM诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。