arXiv:2409.01437cs.CVcs.AI2024-09被引 40

构建首个胃肠镜图文问答数据集,助力医疗AI诊断研究。

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

  • 基于现有数据集扩展,添加多类型问答标注。
  • 含6500张图像,支持是/否、选择、位置、计数等提问。
  • 适用于视觉问答、图像生成等任务,适合医学AI研究者。

我们提出Kvasir-VQA,一个从HyperKvasir和Kvasir-Instrument数据集扩展而来的图文问答数据集,新增问题与答案标注,以支持胃肠镜诊断中的高级机器学习任务。该数据集包含6,500张标注图像,涵盖多种胃肠道疾病和手术器械,支持是/否、选择、位置及数值计数等多种问题类型。可用于图像描述生成、视觉问答(VQA)、基于文本的合成医学图像生成、目标检测与分类等任务。实验表明,该数据集在三个选定任务上训练模型效果显著,验证了其在医学图像分析与诊断中的应用潜力。我们还为每项任务提供了评估指标,凸显数据集的可用性与多功能性。数据集及相关资源可于https://datasets.simula.no/kvasir-vqa获取。

原文摘要 · Abstract (English)

We introduce Kvasir-VQA, an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations to facilitate advanced machine learning tasks in Gastrointestinal (GI) diagnostics. This dataset comprises 6,500 annotated images spanning various GI tract conditions and surgical instruments, and it supports multiple question types including yes/no, choice, location, and numerical count. The dataset is intended for applications such as image captioning, Visual Question Answering (VQA), text-based generation of synthetic medical images, object detection, and classification. Our experiments demonstrate the dataset's effectiveness in training models for three selected tasks, showcasing significant applications in medical image analysis and diagnostics. We also present evaluation metrics for each task, highlighting the usability and versatility of our dataset. The dataset and supporting artifacts are available at https://datasets.simula.no/kvasir-vqa.

医学图像视觉问答数据集胃肠镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。