arXiv:2503.02157cs.CVcs.AI2025-03被引 16

构建医疗视觉语言模型幻觉评估基准,揭示三大幻觉成因及现有缓解方法局限。

MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models

  • 按视觉误判、知识缺失、上下文错位三类成因系统评估幻觉
  • 在11个模型上测试发现现有缓解技术对知识与上下文错误效果有限
  • 适合医疗AI研发者和评测人员,推动更可靠的医学大模型发展

大型视觉语言模型(LVLMs)在医疗领域日益重要,但医学视觉语言模型(Med-LVLMs)常因专业能力不足和医疗应用复杂性产生幻觉。现有基准无法基于幻觉成因进行有效评估,也缺乏对缓解策略的检验。为此,我们提出MedHEval,一个新型基准,通过将幻觉归因于三类根本原因——视觉误判、知识缺失、上下文错位——系统评估其在医学生问答任务中的表现。我们构建了涵盖闭合与开放式问题的多样化医疗视觉问答数据集,并采用全面评估指标。实验覆盖11个主流(医学)视觉语言模型,评估7种先进幻觉缓解技术。结果表明,当前Med-LVLMs在不同成因引发的幻觉上均表现不佳,现有缓解方法整体效果有限,尤其对知识型和上下文相关错误改善甚微。研究强调需加强对齐训练与针对性缓解策略以提升模型可靠性。MedHEval为医疗幻觉的评估与缓解提供了标准化框架,助力更可信医学大模型的发展。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) are becoming increasingly important in the medical domain, yet Medical LVLMs (Med-LVLMs) frequently generate hallucinations due to limited expertise and the complexity of medical applications. Existing benchmarks fail to effectively evaluate hallucinations based on their underlying causes and lack assessments of mitigation strategies. To address this gap, we introduce MedHEval, a novel benchmark that systematically evaluates hallucinations and mitigation strategies in Med-LVLMs by categorizing them into three underlying causes: visual misinterpretation, knowledge deficiency, and context misalignment. We construct a diverse set of close- and open-ended medical VQA datasets with comprehensive evaluation metrics to assess these hallucination types. We conduct extensive experiments across 11 popular (Med)-LVLMs and evaluate 7 state-of-the-art hallucination mitigation techniques. Results reveal that Med-LVLMs struggle with hallucinations arising from different causes while existing mitigation methods show limited effectiveness, especially for knowledge- and context-based errors. These findings underscore the need for improved alignment training and specialized mitigation strategies to enhance Med-LVLMs' reliability. MedHEval establishes a standardized framework for evaluating and mitigating medical hallucinations, guiding the development of more trustworthy Med-LVLMs.

医疗AI幻觉评估视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。