arXiv:2508.10869cs.CVcs.AI2025-08被引 4

构建可解释的胃肠道影像问答系统,提升AI医疗决策可信度。

Medico 2025: Visual Question Answering for Gastrointestinal Imaging

  • 基于胃肠道内镜图像,回答临床相关视觉问题并生成可解释理由。
  • 使用6500张图像与15.9万条复杂问答对,在Kvasir-VQA-x1数据集上评估性能。
  • 适合医学AI、可解释性研究及临床辅助诊断系统开发者关注。

Medico 2025挑战赛聚焦胃肠道(GI)影像的视觉问答(VQA),作为MediaEval任务系列的一部分。该挑战旨在开发可解释人工智能(XAI)模型,根据胃肠道内镜图像回答临床相关问题,并提供符合医学推理逻辑的可解释性说明。比赛包含两个子任务:(1) 利用Kvasir-VQA-x1数据集回答多样化的视觉问题;(2) 生成多模态解释以支持临床决策。Kvasir-VQA-x1数据集由6,500张图像和159,549条复杂问答对构成,是本次挑战的基准。通过量化指标与专家评审的可解释性评估相结合,推动医疗影像分析中可信AI的发展。参赛指南、数据获取及更新说明详见官方仓库:https://github.com/simula/MediaEval-Medico-2025。

原文摘要 · Abstract (English)

The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on developing Explainable Artificial Intelligence (XAI) models that answer clinically relevant questions based on GI endoscopy images while providing interpretable justifications aligned with medical reasoning. It introduces two subtasks: (1) answering diverse types of visual questions using the Kvasir-VQA-x1 dataset, and (2) generating multimodal explanations to support clinical decision-making. The Kvasir-VQA-x1 dataset, created from 6,500 images and 159,549 complex question-answer (QA) pairs, serves as the benchmark for the challenge. By combining quantitative performance metrics and expert-reviewed explainability assessments, this task aims to advance trustworthy Artificial Intelligence (AI) in medical image analysis. Instructions, data access, and an updated guide for participation are available in the official competition repository: https://github.com/simula/MediaEval-Medico-2025

视觉问答医疗AI可解释性胃肠影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。