用模型内部信号实现高效精准的视觉问答幻觉检测
FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering
- 融合解码不确定性和跨模态特征,构建轻量级单阶段检测网络
- 在多个VQA数据集上显著优于现有方法,检测效率提升3倍以上
- 无需人工标注,自动生成监督信号,适合工业级部署
视觉问答中的忠实性幻觉指模型生成流畅但缺乏视觉依据的答案,严重影响其在安全关键场景下的可靠性。现有检测方法分为两类:依赖外部模型或知识库的验证方法,计算开销大且受限于外部资源质量;以及基于不确定性估计的方法,仅捕捉部分不确定性,未能充分挖掘多样失效模式的内部信号。为此,本文提出FaithSCAN:一种轻量级网络,通过融合分词级解码不确定性、中间视觉表示和跨模态对齐特征,利用分支证据编码与不确定性感知注意力进行融合。同时将大模型作为裁判的思想扩展至VQA幻觉检测,提出低成本自动生成模型相关监督信号的方法,实现无须昂贵人工标注的有监督训练,同时保持高检测精度。多基准测试显示,FaithSCAN在有效性和效率上均显著优于现有方法。深入分析表明,幻觉源于视觉感知、跨模态推理和语言解码中的系统性内部状态变化。不同内部信号提供互补诊断线索,且幻觉模式随模型架构变化,为多模态幻觉成因提供了新见解。
原文摘要 · Abstract (English)
Faithfulness hallucinations in VQA occur when vision-language models produce fluent yet visually ungrounded answers, severely undermining their reliability in safety-critical applications. Existing detection methods mainly fall into two categories: external verification approaches relying on auxiliary models or knowledge bases, and uncertainty-driven approaches using repeated sampling or uncertainty estimates. The former suffer from high computational overhead and are limited by external resource quality, while the latter capture only limited facets of model uncertainty and fail to sufficiently explore the rich internal signals associated with the diverse failure modes. Both paradigms thus have inherent limitations in efficiency, robustness, and detection performance. To address these challenges, we propose FaithSCAN: a lightweight network that detects hallucinations by exploiting rich internal signals of VLMs, including token-level decoding uncertainty, intermediate visual representations, and cross-modal alignment features. These signals are fused via branch-wise evidence encoding and uncertainty-aware attention. We also extend the LLM-as-a-Judge paradigm to VQA hallucination and propose a low-cost strategy to automatically generate model-dependent supervision signals, enabling supervised training without costly human labels while maintaining high detection accuracy. Experiments on multiple VQA benchmarks show that FaithSCAN significantly outperforms existing methods in both effectiveness and efficiency. In-depth analysis shows hallucinations arise from systematic internal state variations in visual perception, cross-modal reasoning, and language decoding. Different internal signals provide complementary diagnostic cues, and hallucination patterns vary across VLM architectures, offering new insights into the underlying causes of multimodal hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。