arXiv:2510.06888cs.IRcs.AI2025-10EMNLP被引 5

首个医学多模态检索基准,评估图文结合的医疗信息查找能力。

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

  • 构建跨5大领域16专科的多模态医疗检索数据集
  • 包含超120万文本与16.4万图文查询,覆盖4类任务
  • 适合研究医疗AI、检索系统及多模态模型的开发者

随着检索增强生成(RAG)应用日益广泛,高性能检索模型的重要性不断提升。在医疗领域,结合文本与图像信息的多模态检索模型在问答、跨模态检索和多模态摘要等下游任务中具有显著优势,因为医疗数据常同时包含文本和图像。然而,目前尚无标准基准来评估这些模型在医疗场景下的表现。为此,我们提出M3Retrieve,一个涵盖5个领域、16个医学专科、4种不同任务的多模态医学检索基准,包含超过120万条文本文档和16.4万条多模态查询,所有数据均在合规授权下收集。我们在此基准上评估了主流多模态检索模型,以探究不同医学专科特有的挑战及其对检索性能的影响。通过发布M3Retrieve,我们旨在推动系统化评估,促进模型创新,并加速更强大、可靠的医疗多模态检索系统的研究进程。数据集与基线代码已开源:https://github.com/AkashGhosh/M3Retrieve。

原文摘要 · Abstract (English)

With the increasing use of RetrievalAugmented Generation (RAG), strong retrieval models have become more important than ever. In healthcare, multimodal retrieval models that combine information from both text and images offer major advantages for many downstream tasks such as question answering, cross-modal retrieval, and multimodal summarization, since medical data often includes both formats. However, there is currently no standard benchmark to evaluate how well these models perform in medical settings. To address this gap, we introduce M3Retrieve, a Multimodal Medical Retrieval Benchmark. M3Retrieve, spans 5 domains,16 medical fields, and 4 distinct tasks, with over 1.2 Million text documents and 164K multimodal queries, all collected under approved licenses. We evaluate leading multimodal retrieval models on this benchmark to explore the challenges specific to different medical specialities and to understand their impact on retrieval performance. By releasing M3Retrieve, we aim to enable systematic evaluation, foster model innovation, and accelerate research toward building more capable and reliable multimodal retrieval systems for medical applications. The dataset and the baselines code are available in this github page https://github.com/AkashGhosh/M3Retrieve.

多模态检索医疗AI基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。