提出医学多模态模型定位先于回答的评估与训练框架,减少幻觉。
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
- 设计新评估基准,检测模型是否准确定位病灶区域。
- 在6.7万组医学问答上验证,定位准确率显著提升。
- 适合医疗AI安全、可解释性研究者使用。
医学大模型在医疗数据理解方面表现出色,但常因定位推理不足而产生与证据矛盾的幻觉。本研究揭示当前医学多模态模型的关键缺陷:回答疾病相关问题时,不分析病灶区域,而是依赖语言模式或关注无关图像区域。为此,提出HEAL-MedVQA(基于定位的医学视觉问答幻觉评估)基准,包含两种创新评估协议,用于检测视觉与文本捷径学习;并构建了67,000个带医生标注病灶分割掩码的VQA对。为提升视觉推理能力,提出“定位后回答”(LobA)框架,通过训练模型先定位目标区域,并自提示强调病灶区域,生成更可靠答案。实验表明,该方法在具有挑战性的HEAL-MedVQA基准上显著优于现有生物医学多模态模型,提升了医学视觉问答的鲁棒性。
原文摘要 · Abstract (English)
Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。