arXiv:2501.13277cs.CV2025-01被引 1

用临床数据辅助CT图像预训练,提升多癌种分析能力

MEDFORM: A Foundation Model for Contrastive Learning of CT Imaging and Clinical Numeric Data in Multi-Cancer Analysis

  • 用临床数值引导CT图像对比学习,实现跨模态协同预训练
  • 在肺癌、乳腺癌、结直肠癌数据上提升分类性能,少样本下仍稳定
  • 适合医疗多模态模型研究者,尤其关注医学基础模型构建

计算机断层扫描(CT)和临床数值数据是癌症评估的关键模态,但因多切片CT结构复杂且专家标注成本高,构建大规模多模态训练数据集仍具挑战。本文提出MEDFORM,一种基于临床数据引导的多模态预训练策略,用于医学基础模型开发。该方法通过多重实例学习(MIL)高效处理CT切片,并采用双阶段预训练:先使用SimCLR进行自监督学习训练CT特征提取器,再通过跨模态对比学习对齐CT与临床数据。模型在三种癌种上预训练:肺癌(141,171切片)、乳腺癌(8,100切片)、结直肠癌(10,393切片)。实验表明,该双阶段预训练策略显著提升癌症分类性能,并在小样本场景中保持鲁棒性。代码已公开于https://github.com/DigitalHealthcareLab/25MultiModalFoundationModel.git。

原文摘要 · Abstract (English)

Computed tomography (CT) and clinical numeric data are essential modalities for cancer evaluation, but building large-scale multimodal training datasets for developing medical foundation models remains challenging due to the structural complexity of multi-slice CT data and high cost of expert annotation. In this study, we propose MEDFORM, a multimodal pre-training strategy that guides CT image representation learning using complementary information from clinical data for medical foundation model development. MEDFORM efficiently processes CT slice through multiple instance learning (MIL) and adopts a dual pre-training strategy: first pretraining the CT slice feature extractor using SimCLR-based self-supervised learning, then aligning CT and clinical modalities through cross-modal contrastive learning. Our model was pre-trained on three different cancer types: lung cancer (141,171 slices), breast cancer (8,100 slices), colorectal cancer (10,393 slices). The experimental results demonstrated that this dual pre-training strategy improves cancer classification performance and maintains robust performance in few-shot learning scenarios. Code available at https://github.com/DigitalHealthcareLab/25MultiModalFoundationModel.git

医学多模态对比学习基础模型影像分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。