首个面向医学视频的多语言多模态多跳推理基准,挑战AI深度理解能力。
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
- 设计多跳推理任务,需跨文本与视觉信息关联定位答案
- 模型在复杂问题上表现远低于人类专家,差距显著
- 适用于医疗AI、多模态理解研究者,推动领域深化
随着多模态理解技术的快速发展,视频理解有望赋能医学教育等专业领域。然而现有基准存在两大局限:(1) 语言单一性——主要局限于英语,缺乏多语言资源;(2) 推理浅层化——问题多为表层信息检索,无法评估深层多模态融合能力。为此,我们提出M3-Med,首个面向医学教学视频理解的多语言、多模态、多跳推理基准。M3-Med包含由医学专家标注的医学问题与对应视频片段,核心创新在于多跳推理任务:模型需先从文本中定位关键实体,再在视频中寻找视觉证据,最后融合双模态信息得出答案。该设计超越简单文本匹配,对模型的深层跨模态理解能力构成重大挑战。我们定义两项任务:单视频时间答案定位(TAGSV)和视频语料库时间答案定位(TAGVC)。我们在M3-Med上评估多个先进模型与大语言模型,结果揭示所有模型与人类专家之间存在显著性能差距,尤其在复杂多跳问题上模型表现急剧下降。M3-Med有效揭示了当前AI模型在专业领域深度跨模态推理中的局限,并为未来研究提供新方向。
原文摘要 · Abstract (English)
With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional domains such as medical education. However, existing benchmarks suffer from two primary limitations: (1) Linguistic Singularity: they are largely confined to English, neglecting the need for multilingual resources; and (2) Shallow Reasoning: their questions are often designed for surface-level information retrieval, failing to properly assess deep multi-modal integration. To address these limitations, we present M3-Med, the first benchmark for Multi-lingual, Multi-modal, and Multi-hop reasoning in Medical instructional video understanding. M3-Med consists of medical questions paired with corresponding video segments, annotated by a team of medical experts. A key innovation of M3-Med is its multi-hop reasoning task, which requires a model to first locate a key entity in the text, then find corresponding visual evidence in the video, and finally synthesize information across both modalities to derive the answer. This design moves beyond simple text matching and poses a substantial challenge to a model's deep cross-modal understanding capabilities. We define two tasks: Temporal Answer Grounding in Single Video (TAGSV) and Temporal Answer Grounding in Video Corpus (TAGVC). We evaluated several state-of-the-art models and Large Language Models (LLMs) on M3-Med. The results reveal a significant performance gap between all models and human experts, especially on the complex multi-hop questions where model performance drops sharply. M3-Med effectively highlights the current limitations of AI models in deep cross-modal reasoning within specialized domains and provides a new direction for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。