arXiv:2411.16778cs.CVcs.AI2024-11ICCV被引 34

构建首个可定位、可解释的胸部X光VQA基准,助力临床AI可信赖诊断。

GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis

  • 引入多模态解释机制,提供答案的视觉与文本依据。
  • 涵盖151,025张图像和160万+问题,是当前最大胸部X光VQA数据集。
  • 支持四种问题类型,更贴近真实临床场景,适合医疗AI研究者使用。

医学视觉问答(Med-VQA)融合计算机视觉与自然语言处理,自动回答医学影像相关的临床问题。然而现有数据集存在两大局限:(1)缺乏对答案的视觉与文本解释,不利于患者及低年资医生理解;(2)问题形式单一,难以反映实际诊疗中多样化需求。为应对这些挑战,我们提出一个大规模、可定位、可解释的胸部X光诊断医学VQA基准GEMeX,包含三项创新:(1)多模态可解释机制,为每个问答对提供详细视觉与文本解释,提升答案可理解性;(2)设计四类问题形式——开放式、封闭式、单选题、多选题,更贴合临床实际。该数据集包含151,025张图像与1,605,575个问题,是目前最大的胸部X光VQA数据集。在12个代表性大视觉语言模型上评估发现其表现不佳,凸显数据集复杂性。我们通过在GEMeX训练集上微调现有模型,提出一个强基线模型,性能显著提升,验证了数据集的有效性。基准数据已开放:https://www.med-vqa.com/GEMeX。

原文摘要 · Abstract (English)

Medical Visual Question Answering (Med-VQA) combines computer vision and natural language processing to automatically answer clinical inquiries about medical images. However, current Med-VQA datasets exhibit two significant limitations: (1) they often lack visual and textual explanations for answers, hindering comprehension for patients and junior doctors; (2) they typically offer a narrow range of question formats, inadequately reflecting the diverse requirements in practical scenarios. These limitations pose significant challenges to the development of a reliable and user-friendly Med-VQA system. To address these challenges, we introduce a large-scale, Groundable, and Explainable Medical VQA benchmark for chest X-ray diagnosis (GEMeX), featuring several innovative components: (1) a multi-modal explainability mechanism that offers detailed visual and textual explanations for each question-answer pair, thereby enhancing answer comprehensibility; (2) four question types, open-ended, closed-ended, single-choice, and multiple-choice, to better reflect practical needs. With 151,025 images and 1,605,575 questions, GEMeX is the currently largest chest X-ray VQA dataset. Evaluation of 12 representative large vision language models (LVLMs) on GEMeX reveals suboptimal performance, underscoring the dataset's complexity. Meanwhile, we propose a strong model by fine-tuning an existing LVLM on the GEMeX training set. The substantial performance improvement showcases the dataset's effectiveness. The benchmark is available at https://www.med-vqa.com/GEMeX.

医学AI视觉问答可解释性胸部X光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。