构建3D医学视觉基础模型,提升病灶诊断与报告生成能力
E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model
- 用自监督学习构建3D CT视觉基础模型,解决数据少、维度高难题
- 引入3D空间卷积压缩特征,保持空间信息同时降低计算量
- 基于BIMCV-R和CT-RATE构建指令数据集,适用于临床场景
3D医学视觉语言模型在疾病诊断与治疗中潜力巨大。然而,相较于2D医学图像,3D医学图像(如CT扫描)面临训练数据有限、维度高等挑战,严重制约了3D医学视觉语言模型的发展。为此,我们收集大量未标注3D CT数据,采用自监督学习构建3D视觉基础模型,用于提取3D视觉特征。通过3D空间卷积聚合并投影高层图像特征,在降低计算复杂度的同时保留空间信息。此外,我们基于BIMCV-R和CT-RATE构建两个指令微调数据集,对3D视觉语言模型进行微调。实验表明,该模型在报告生成、视觉问答和疾病诊断任务上均优于现有方法。代码与数据将尽快公开。
原文摘要 · Abstract (English)
The development of 3D medical vision-language models holds significant potential for disease diagnosis and patient treatment. However, compared to 2D medical images, 3D medical images, such as CT scans, face challenges related to limited training data and high dimension, which severely restrict the progress of 3D medical vision-language models. To address these issues, we collect a large amount of unlabeled 3D CT data and utilize self-supervised learning to construct a 3D visual foundation model for extracting 3D visual features. Then, we apply 3D spatial convolutions to aggregate and project high-level image features, reducing computational complexity while preserving spatial information. We also construct two instruction-tuning datasets based on BIMCV-R and CT-RATE to fine-tune the 3D vision-language model. Our model demonstrates superior performance compared to existing methods in report generation, visual question answering, and disease diagnosis. Code and data will be made publicly available soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。