arXiv:2506.17939cs.CVcs.AI2025-06中稿 · ACM MM 2025被引 5

构建可解释的医学视觉问答数据集,提升模型推理可信度

GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning

  • 设计区域感知的多模态思维链数据集,每步推理对应图像特定区域
  • 用可验证奖励机制训练模型,使推理过程与答案高度对齐
  • 仅需1/8数据量即达相近性能,适合医疗AI可解释性研究

医学视觉问答旨在通过医疗影像回答自然语言问题以辅助临床决策。尽管多模态学习取得进展,现有方法仍存在答案不可靠、解释性差的问题,影响医生和患者对模型输出的信任。为此,本文提出一个区域感知的多模态思维链(RMCoT)数据集,其答案生成前包含一系列中间推理步骤,明确关联医学图像中的相关视觉区域,实现细粒度可解释性。此外,引入一种新颖的可验证奖励机制用于强化学习后训练,引导模型推理过程与最终答案对齐。显著的是,该方法仅使用1/8的训练数据即可达到相当性能,证明了该方案的高效性与有效性。数据集已开放获取:https://www.med-vqa.com/GEMeX/

原文摘要 · Abstract (English)

Medical visual question answering aims to support clinical decision-making by enabling models to answer natural language questions based on medical images. While recent advances in multi-modal learning have significantly improved performance, current methods still suffer from limited answer reliability and poor interpretability, impairing the ability of clinicians and patients to understand and trust model outputs. To address these limitations, this work first proposes a Region-Aware Multimodal Chain-of-Thought (RMCoT) dataset, in which the process of producing an answer is preceded by a sequence of intermediate reasoning steps that explicitly ground relevant visual regions of the medical image, thereby providing fine-grained explainability. Furthermore, we introduce a novel verifiable reward mechanism for reinforcement learning to guide post-training, improving the alignment between the model's reasoning process and its final answer. Remarkably, our method achieves comparable performance using only one-eighth of the training data, demonstrating the efficiency and effectiveness of the proposal. The dataset is available at https://www.med-vqa.com/GEMeX/.

医学视觉问答可解释AI多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。