arXiv:2603.19516cs.CVcs.AI2026-03被引 2

构建胃癌多模态数据集,推动视觉语言模型临床应用

Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis

  • 构建1700例胃癌多模态数据,包含影像、生化指标与诊断报告
  • 在5项临床任务中评估模型表现,揭示其跨模态关联能力
  • 适合医学AI研究者和多模态模型开发者使用

近期视觉语言模型(VLMs)在自然领域展现出强大的泛化与多模态推理能力,但在医疗诊断中的应用受限于缺乏全面且结构化的数据集。为推进VLM在临床场景的发展,特别是胃癌分析,我们提出Gastric-X,一个大规模多模态基准数据集,包含1.7K例胃癌病例。每例包含配对的静息与动态CT扫描、内镜图像、一组结构化生化指标、专家撰写的诊断报告以及肿瘤区域的边界框标注,真实反映临床流程。我们系统评估了当前VLMs在五项核心任务中的表现:视觉问答(VQA)、报告生成、跨模态检索、疾病分类与病灶定位。这些任务模拟临床工作流中的关键环节,从视觉理解到多模态决策支持。通过评估,不仅考察模型性能,更探究其理解本质:当前VLM能否有意义地关联生化信号、空间肿瘤特征与文本报告?我们期望Gastric-X能促进机器智能与医生认知及证据推理过程的对齐,并启发下一代医疗VLM的发展。

原文摘要 · Abstract (English)

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured datasets that capture real clinical workflows. To advance the development of VLMs for clinical applications, particularly in gastric cancer, we introduce Gastric-X, a large-scale multimodal benchmark for gastric cancer analysis providing 1.7K cases. Each case in Gastric-X includes paired resting and dynamic CT scans, endoscopic image, a set of structured biochemical indicators, expert-authored diagnostic notes, and bounding box annotations of tumor regions, reflecting realistic clinical conditions. We systematically examine the capability of recent VLMs on five core tasks: Visual Question Answering (VQA), report generation, cross-modal retrieval, disease classification, and lesion localization. These tasks simulate critical stages of clinical workflow, from visual understanding and reasoning to multimodal decision support. Through this evaluation, we aim not only to assess model performance but also to probe the nature of VLM understanding: Can current VLMs meaningfully correlate biochemical signals with spatial tumor features and textual reports? We envision Gastric-X as a step toward aligning machine intelligence with the cognitive and evidential reasoning processes of physicians, and as a resource to inspire the development of next-generation medical VLMs.

多模态胃癌视觉语言模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。