arXiv:2412.11056cs.CV2024-12综述被引 6

构建医疗视频问答新任务,推动AI理解医学视频并生成操作指导。

Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track

  • 设计医疗视频问答与指令生成任务,融合语言与视觉模态。
  • 强调医学视频在急救、教育等场景中提供直观答案的优势。
  • 适合关注医疗AI、多模态理解与临床辅助系统的研究者。

人工智能的重要目标之一是构建能通过自然语言查询与视觉世界(图像和视频)交互的多模态系统。以往医疗问答研究主要聚焦文本和视觉(图像)模态,对需演示才能回答的问题效率较低。近年来,得益于大规模语言-视觉数据集和高效深度神经网络的发展,视觉-语言任务如视觉描述生成、视觉问答和自然语言视频定位取得了显著进展。现有工作多集中于开放域应用的数据集构建与解决方案。我们认为,医学视频能为急救、医疗紧急情况和医学教育类问题提供最准确的答案。随着人工智能在支持临床决策和提升患者参与度方面的关注度上升,亟需探索此类挑战并开发高效的医疗语言-视频理解与生成算法。为此,我们引入新任务,旨在促进设计能够理解医学视频、以视觉方式回答自然语言问题,并具备从医学视频生成操作步骤的多模态能力的系统。这些任务有望推动下游复杂应用的发展,惠及公众与医疗从业者。

原文摘要 · Abstract (English)

One of the key goals of artificial intelligence (AI) is the development of a multimodal system that facilitates communication with the visual world (image and video) using a natural language query. Earlier works on medical question answering primarily focused on textual and visual (image) modalities, which may be inefficient in answering questions requiring demonstration. In recent years, significant progress has been achieved due to the introduction of large-scale language-vision datasets and the development of efficient deep neural techniques that bridge the gap between language and visual understanding. Improvements have been made in numerous vision-and-language tasks, such as visual captioning visual question answering, and natural language video localization. Most of the existing work on language vision focused on creating datasets and developing solutions for open-domain applications. We believe medical videos may provide the best possible answers to many first aid, medical emergency, and medical education questions. With increasing interest in AI to support clinical decision-making and improve patient engagement, there is a need to explore such challenges and develop efficient algorithms for medical language-video understanding and generation. Toward this, we introduced new tasks to foster research toward designing systems that can understand medical videos to provide visual answers to natural language questions, and are equipped with multimodal capability to generate instruction steps from the medical video. These tasks have the potential to support the development of sophisticated downstream applications that can benefit the public and medical professionals.

医疗AI多模态视频问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。