聚焦医疗教学视频的多模态多语言多跳问答挑战,推动智能医疗系统发展。
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
- 构建三赛道:单视频时序定位、视频语料检索、跨视频时序定位
- 支持多语言查询与跨模态推理,需解决复杂医学问题链
- 适合医疗AI、多模态系统研究者,助力多语言医疗教育
继2023年佛山首届CMIVQA和2024年杭州第二届MMIVQA之后,2025年NLPCC引入全新任务M4IVQA,聚焦多模态、多语言、多跳医学教学视频问答。该挑战评估模型在医学教学视频中融合视觉、语言信息,理解多语言提问,并回答需跨模态推理的复杂问题的能力。任务包含三个赛道:单视频时序答案定位(M4TAGSV)、多模态多语言视频语料检索(M4VCR)以及跨视频时序答案定位(M4TAGVC)。参赛者需开发能处理视频与文本数据、理解多语言查询并给出精准答案的算法。本挑战旨在推动医疗场景下多模态推理系统创新,促进智能应急响应与多语言医学教育平台发展。官方页面:https://cmivqa.github.io/
原文摘要 · Abstract (English)
Following the successful hosts of the 1-st (NLPCC 2023 Foshan) CMIVQA and the 2-rd (NLPCC 2024 Hangzhou) MMIVQA challenges, this year, a new task has been introduced to further advance research in multi-modal, multilingual, and multi-hop medical instructional question answering (M4IVQA) systems, with a specific focus on medical instructional videos. The M4IVQA challenge focuses on evaluating models that integrate information from medical instructional videos, understand multiple languages, and answer multi-hop questions requiring reasoning over various modalities. This task consists of three tracks: multi-modal, multilingual, and multi-hop Temporal Answer Grounding in Single Video (M4TAGSV), multi-modal, multilingual, and multi-hop Video Corpus Retrieval (M4VCR) and multi-modal, multilingual, and multi-hop Temporal Answer Grounding in Video Corpus (M4TAGVC). Participants in M4IVQA are expected to develop algorithms capable of processing both video and text data, understanding multilingual queries, and providing relevant answers to multi-hop medical questions. We believe the newly introduced M4IVQA challenge will drive innovations in multimodal reasoning systems for healthcare scenarios, ultimately contributing to smarter emergency response systems and more effective medical education platforms in multilingual communities. Our official website is https://cmivqa.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。