arXiv:2411.14522cs.CV2024-11被引 24

构建医疗多模态数据集与模型,提升AI在医学诊断中的表现

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

  • 用上百个专业医疗数据集转换成高质量图文对,形成大体量医疗多模态数据集
  • 基于新数据集训练三阶段模型,显著提升视觉与文本信息融合能力
  • 在多种医学图像任务中达到顶尖水平,适合临床辅助决策研究者使用

尽管通用人工智能取得显著进展,但在医疗领域仍受限于缺乏专业医学知识。为此,我们构建了GMAI-VL-5.5M,一个通过将数百个具有不同标注的专科医疗数据集转换为高质量图像-文本对形成的多模态医学数据集。该数据集涵盖全面的任务类型、多样模态和丰富的图文数据。在此基础上,我们开发了GMAI-VL,一种通用医学视觉语言模型,采用三阶段训练策略,增强了视觉与文本信息的融合能力。该方法显著提升了模型处理多模态数据的能力,支持更精准的诊断与临床决策。实验表明,GMAI-VL在多项多模态医学任务中(包括视觉问答和医学图像诊断)均达到当前最优性能。

原文摘要 · Abstract (English)

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis.

多模态医学AI视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。