arXiv:2503.20504cs.CV2025-03被引 5

提出统一框架,用视觉引导的语义熵检测医疗多模态模型幻觉。

UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models

  • 通过对比原图与扭曲图的语义分布,增强视觉对不确定性的约束。
  • 在六组医学数据集上,检测准确率显著优于现有方法。
  • 适合关注医疗AI可信度、幻觉检测的研究者与开发者。

视觉语言模型(VLMs)在医学图像理解中潜力巨大,尤其适用于视觉报告生成(VRG)和视觉问答(VQA),但可能生成与视觉证据矛盾的幻觉响应,限制临床应用。尽管基于不确定性的幻觉检测方法直观有效,但在医学VLM中受限:语义熵(SE)在纯文本大模型中有效,但在医学VLM中因强语言先验导致过度自信而可靠性下降。为此,本文提出UniVRSE——一种面向医学VLM的统一视觉条件化响应语义熵框架。该方法通过对比原始图像-文本对与视觉扭曲后的对应对的语义预测分布,提升不确定性估计中的视觉引导作用;熵值越高,幻觉风险越大。针对VQA,基于图像-问题对进行评估;针对VRG,将报告分解为声明,生成验证问题,并在声明层面应用视觉条件化熵估计。为评估幻觉检测效果,提出统一流程:在医学数据集生成响应,并通过事实一致性评估生成幻觉标签。现有方法依赖主观标准或模态特定规则,为提升可靠性,引入原子事实对齐比(ALFA),量化细粒度事实一致性。基于ALFA的标签提供可靠基准。在六个医学VQA/VRG数据集和三个VLM上实验表明,UniVRSE显著优于现有方法,具备强跨模态泛化能力。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have great potential for medical image understanding, particularly in Visual Report Generation (VRG) and Visual Question Answering (VQA), but they may generate hallucinated responses that contradict visual evidence, limiting clinical deployment. Although uncertainty-based hallucination detection methods are intuitive and effective, they are limited in medical VLMs. Specifically, Semantic Entropy (SE), effective in text-only LLMs, becomes less reliable in medical VLMs due to their overconfidence from strong language priors. To address this challenge, we propose UniVRSE, a Unified Vision-conditioned Response Semantic Entropy framework for hallucination detection in medical VLMs. UniVRSE strengthens visual guidance during uncertainty estimation by contrasting the semantic predictive distributions derived from an original image-text pair and a visually distorted counterpart, with higher entropy indicating hallucination risk. For VQA, UniVRSE works on the image-question pair, while for VRG, it decomposes the report into claims, generates verification questions, and applies vision-conditioned entropy estimation at the claim level. To evaluate hallucination detection, we propose a unified pipeline that generates responses on medical datasets and derives hallucination labels via factual consistency assessment. However, current evaluation methods rely on subjective criteria or modality-specific rules. To improve reliability, we introduce Alignment Ratio of Atomic Facts (ALFA), a novel method that quantifies fine-grained factual consistency. ALFA-derived labels provide ground truth for robust benchmarking. Experiments on six medical VQA/VRG datasets and three VLMs show UniVRSE significantly outperforms existing methods with strong cross-modal generalization.

幻觉检测医疗AI视觉语言模型不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。