新任务评估医疗视频问答中不同难度问题的多模态理解能力。
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

- 按证据复杂度区分问题类型,分文本、视觉和跨模态三类难度。
- 覆盖急救、护理等场景,含人工标注的难度分级数据集。
- 适合医疗AI、多模态推理与可解释性研究者参考。
继NLPCC 2023–2025年的CMIVQA、MMI-VQA和M4IVQA挑战之后,我们推出NLPCC 2026的难度感知医学教学视频问答(DA-MIVQA)共享任务。该任务在多语言、多模态医学视频基准上扩展,明确区分需不同证据类型与复杂度的问题:简单问题可仅从字幕文本回答,复杂问题则需视觉定位、流程理解与跨模态融合。挑战包含三个赛道:单视频时序答案定位(DA-TAGSV)、视频语料库检索(DA-VCR)、视频语料库时序答案定位(DA-TAGVC)。数据集源自公开医疗教学频道,涵盖急救、应急响应、康复、护理及通用医学教育等多样化场景,经人工验证并标注难度等级。本文介绍任务动机、数据构建、评估协议、参赛情况、竞赛结果及代表性系统。DA-MIVQA为评估医疗教学视频问答系统在不同文本、视觉、时间与流程推理需求下的表现提供了实用基准。
原文摘要 · Abstract (English)
Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。