arXiv:2512.01922cs.CV2025-12被引 11

医学视觉语言模型防幻觉新方法,不降速还能提准确率

Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding

论文配图:Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding
图 1 · 摘自论文原文
  • 通过动态筛选视觉相关词汇,精简冗余信息提升可靠性
  • 在8个医学数据集上平均提升事实准确性13%,幻觉率降低6%
  • 无需额外推理步骤,适合医疗影像问答与报告生成场景

大型视觉语言模型(LVLMs)正广泛应用于医学视觉问答和影像报告生成等任务,但其仍易产生看似合理却错误的幻觉输出。现有自然图像领域的缓解方法多依赖二次解码或回滚机制,显著增加推理延迟,且常具领域局限性,可能引入模态间或生成内容与真实内容之间的错位。本文提出Med-VCD,一种稀疏视觉对比解码方法,可在不增加推理时间的前提下有效抑制医学LVLM中的幻觉。该方法采用新型令牌稀疏化策略,实时选择具有视觉依据的词元,剔除冗余信息的同时保留关键视觉上下文,在效率与可靠性间取得平衡。在涵盖眼科学、放射学和病理科的8个医学数据集上的评估显示,Med-VCD使事实准确率平均提升13%,幻觉准确率相对基线提高6%。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) are now central to healthcare applications such as medical visual question answering and imaging report generation. Yet, these models remain vulnerable to hallucination outputs that appear plausible but are in fact incorrect. In the natural image domain, several decoding strategies have been proposed to mitigate hallucinations by reinforcing visual evidence, but most rely on secondary decoding or rollback procedures that substantially slow inference. Moreover, existing solutions are often domain-specific and may introduce misalignment between modalities or between generated and ground-truth content. We introduce Med-VCD, a sparse visual-contrastive decoding method that mitigates hallucinations in medical LVLMs without the time overhead of secondary decoding. Med-VCD incorporates a novel token-sparsification strategy that selects visually informed tokens on the fly, trimming redundancy while retaining critical visual context and thus balancing efficiency with reliability. Evaluations on eight medical datasets, spanning ophthalmology, radiology, and pathology tasks in visual question answering, report generation, and dedicated hallucination benchmarks, show that Med-VCD raises factual accuracy by an average of 13\% and improves hallucination accuracy by 6\% relative to baseline medical LVLMs.

医学AI幻觉抑制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。