arXiv:2506.19513cs.CVcs.LG2025-06

用证据冲突检测大模型视觉幻觉,提升安全关键场景可靠性

Visual hallucination detection in large vision-language models via evidential conflict

  • 基于证据理论构建不确定性检测机制,捕捉高层特征冲突
  • 在关系推理任务中发现更严重幻觉,三模型平均准确率提升7%以上
  • 首个结合感知与推理的评测数据集,适合模型安全评估研究者

尽管大型视觉语言模型(LVLMs)具备出色的多模态能力,但视觉输入与文本输出之间常出现不一致——我们称之为视觉幻觉。这一关键可靠性问题在安全敏感的人工智能应用中带来重大风险,亟需全面的评估基准和有效检测方法。我们发现,现有以视觉为中心的幻觉评测主要从感知角度评估,忽视了高级推理引发的幻觉。为此,我们构建了感知-推理评估幻觉(PRE-HAL)数据集,可系统评估LVLM在实例、场景和关系等多重视觉语义下的感知与推理能力。在该新基准上的全面评估揭示了更多视觉漏洞,尤其在更具挑战性的关系推理任务中。为应对此问题,我们首次提出基于德普斯特-谢弗理论(DST)的视觉幻觉检测方法,通过不确定性估计捕获模型推理阶段的高阶特征冲突。该方法采用简单的质量函数,降低幂集上证据融合的计算复杂度。我们在LLaVA-v1.5、mPLUG-Owl2和mPLUG-Owl3三款前沿模型上进行了广泛评估,结果表明该方法优于五种基线不确定性指标,在三个模型上分别实现4%、10%和7%的平均AUROC提升。代码已开源。

原文摘要 · Abstract (English)

Despite the remarkable multimodal capabilities of Large Vision-Language Models (LVLMs), discrepancies often occur between visual inputs and textual outputs--a phenomenon we term visual hallucination. This critical reliability gap poses substantial risks in safety-critical Artificial Intelligence (AI) applications, necessitating a comprehensive evaluation benchmark and effective detection methods. Firstly, we observe that existing visual-centric hallucination benchmarks mainly assess LVLMs from a perception perspective, overlooking hallucinations arising from advanced reasoning capabilities. We develop the Perception-Reasoning Evaluation Hallucination (PRE-HAL) dataset, which enables the systematic evaluation of both perception and reasoning capabilities of LVLMs across multiple visual semantics, such as instances, scenes, and relations. Comprehensive evaluation with this new benchmark exposed more visual vulnerabilities, particularly in the more challenging task of relation reasoning. To address this issue, we propose, to the best of our knowledge, the first Dempster-Shafer theory (DST)-based visual hallucination detection method for LVLMs through uncertainty estimation. This method aims to efficiently capture the degree of conflict in high-level features at the model inference phase. Specifically, our approach employs simple mass functions to mitigate the computational complexity of evidence combination on power sets. We conduct an extensive evaluation of state-of-the-art LVLMs, LLaVA-v1.5, mPLUG-Owl2 and mPLUG-Owl3, with the new PRE-HAL benchmark. Experimental results indicate that our method outperforms five baseline uncertainty metrics, achieving average AUROC improvements of 4%, 10%, and 7% across three LVLMs. Our code is available at https://github.com/HT86159/Evidential-Conflict.

视觉幻觉大模型安全证据理论评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。