arXiv:2506.22900cs.CVcs.CL2025-06被引 10

MOTOR通过视觉与文本联合重排序,提升医学影像问答的准确性。

MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering

  • 利用视觉与文本双重信息进行检索重排序
  • 在多个数据集上平均准确率提升6.45%
  • 适合医疗AI研发与临床辅助系统开发者

医学视觉问答(MedVQA)通过图像相关问题提供上下文丰富的回答,在临床决策中至关重要。尽管视觉语言模型(VLMs)广泛应用,但常生成事实错误答案。检索增强生成通过引入外部信息缓解此问题,但易检索无关内容,削弱模型推理能力。现有方法虽通过查询-文本对齐重排序提升相关性,却忽视了对医学诊断至关重要的视觉或多模态上下文。本文提出MOTOR,一种基于接地描述和最优传输的新型多模态检索与重排序方法。该方法结合文本与视觉信息,捕捉查询与检索内容间的深层关联,从而识别更符合临床实际的上下文以增强VLM输入。实证分析与专家评估表明,MOTOR在多个MedVQA数据集上表现优异,平均准确率超越现有最佳方法6.45%。代码已公开于https://github.com/BioMedIA-MBZUAI/MOTOR。

原文摘要 · Abstract (English)

Medical visual question answering (MedVQA) plays a vital role in clinical decision-making by providing contextually rich answers to image-based queries. Although vision-language models (VLMs) are widely used for this task, they often generate factually incorrect answers. Retrieval-augmented generation addresses this challenge by providing information from external sources, but risks retrieving irrelevant context, which can degrade the reasoning capabilities of VLMs. Re-ranking retrievals, as introduced in existing approaches, enhances retrieval relevance by focusing on query-text alignment. However, these approaches neglect the visual or multimodal context, which is particularly crucial for medical diagnosis. We propose MOTOR, a novel multimodal retrieval and re-ranking approach that leverages grounded captions and optimal transport. It captures the underlying relationships between the query and the retrieved context based on textual and visual information. Consequently, our approach identifies more clinically relevant contexts to augment the VLM input. Empirical analysis and human expert evaluation demonstrate that MOTOR achieves higher accuracy on MedVQA datasets, outperforming state-of-the-art methods by an average of 6.45%. Code is available at https://github.com/BioMedIA-MBZUAI/MOTOR.

医学问答多模态检索增强最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。