arXiv:2506.09958cs.CVcs.LG2025-06被引 18

构建胃镜医学视觉问答数据集,提升AI临床推理与鲁棒性评估能力

Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy

  • 用大模型生成15.9万条分层复杂度的问答回答,模拟真实临床推理
  • 引入多种成像伪影增强,使模型在真实场景下测试更可靠
  • 支持标准问答与抗干扰鲁棒性双评测,适合临床级多模态研究

医学视觉问答(MedVQA)是发展临床辅助决策系统的重要方向,但现有数据集常缺乏临床复杂性和视觉多样性。为填补这一空白,我们推出了针对胃肠道内镜的Kvasir-VQA-x1数据集。该数据集在原始Kvasir-VQA基础上新增159,549个问题-答案对,专为检验深度临床推理能力而设计。我们采用大语言模型构建系统化生成方法,并按复杂度分层以更好评估模型推理能力。为确保模型适应真实临床环境,我们引入多种模拟成像伪影的视觉增强。数据集支持两个评估赛道:标准VQA性能与对视觉扰动的鲁棒性测试。通过提供更具挑战性和临床相关性的基准,Kvasir-VQA-x1旨在加速可靠、高效的多模态AI系统在临床中的应用。数据集完全开放并遵循FAIR原则,代码与数据可在GitHub及Hugging Face获取。

原文摘要 · Abstract (English)

Medical Visual Question Answering (MedVQA) is a promising field for developing clinical decision support systems, yet progress is often limited by the available datasets, which can lack clinical complexity and visual diversity. To address these gaps, we introduce Kvasir-VQA-x1, a new, large-scale dataset for gastrointestinal (GI) endoscopy. Our work significantly expands upon the original Kvasir-VQA by incorporating 159,549 new question-answer pairs that are designed to test deeper clinical reasoning. We developed a systematic method using large language models to generate these questions, which are stratified by complexity to better assess a model's inference capabilities. To ensure our dataset prepares models for real-world clinical scenarios, we have also introduced a variety of visual augmentations that mimic common imaging artifacts. The dataset is structured to support two main evaluation tracks: one for standard VQA performance and another to test model robustness against these visual perturbations. By providing a more challenging and clinically relevant benchmark, Kvasir-VQA-x1 aims to accelerate the development of more reliable and effective multimodal AI systems for use in clinical settings. The dataset is fully accessible and adheres to FAIR data principles, making it a valuable resource for the wider research community. Code and data: https://github.com/Simula/Kvasir-VQA-x1 and https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1

医学视觉问答多模态胃镜数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。