arXiv:2412.05876cs.CVcs.AI2024-12被引 13

用多粒度医学知识增强3D影像与报告的预训练,提升模型泛化能力。

MG-3D: Multi-Grained Knowledge-Enhanced 3D Medical Vision-Language Pre-training

  • 通过跨模态对齐与局部重建,统一患者内部影像与报告语义
  • 基于细粒度报告关联建模患者间视觉语义关系,提升特征区分性
  • 在47.1K数据上预训练,适用于多种临床任务,适合医疗AI研究者

3D医学影像分析在临床中至关重要,但标注数据稀缺与泛化能力不足限制了AI模型发展。放射科报告易获取,可作为弱监督信号。然而,3D医学影像中的大规模视觉-语言预训练仍不充分,尤其对患者间多粒度报告语义及其关联的研究不足,导致大规模体数据与报告数据未被充分利用。针对患者内跨模态语义一致性与患者间语义相关性,我们提出MG-3D方法,在47.1K规模数据上进行多任务视觉-语言预训练:1)通过跨模态全局对齐与互补模态引导的局部重建,建立体积语义与患者多粒度医学知识的对应关系,确保不同模态特征协同表达同一语义;2)基于细粒度报告关联建模患者间视觉语义关系,并通过对比学习保持对个体差异的敏感性,增强特征判别力。进一步探究缩放规律以评估性能提升潜力。在九项单模态与跨模态临床任务上进行全面评估,内外部数据实验均验证了MG-3D出色的可迁移性、可扩展性与泛化能力,展现出推动3D医学影像特征表示发展的潜力。代码将公开于:https://github.com/Xuefeng-Ni/MG-3D。

原文摘要 · Abstract (English)

3D medical image analysis is pivotal in numerous clinical applications. However, the scarcity of labeled data and limited generalization capabilities hinder the advancement of AI-empowered models. Radiology reports are easily accessible and can serve as weakly-supervised signals. However, large-scale vision-language pre-training (VLP) remains underexplored in 3D medical image analysis. Specifically, the insufficient investigation into multi-grained radiology semantics and their correlations across patients leads to underutilization of large-scale volume-report data. Considering intra-patient cross-modal semantic consistency and inter-patient semantic correlations, we propose a multi-task VLP method, MG-3D, pre-trained on large-scale data (47.1K), addressing the challenges by the following two aspects: 1) Establishing the correspondence between volume semantics and multi-grained medical knowledge of each patient with cross-modal global alignment and complementary modality-guided local reconstruction, ensuring intra-patient features of different modalities cohesively represent the same semantic content; 2) Correlating inter-patient visual semantics based on fine-grained report correlations across patients, and keeping sensitivity to global individual differences via contrastive learning, enhancing the discriminative feature representation. Furthermore, we delve into the scaling law to explore potential performance improvements. Comprehensive evaluations across nine uni- and cross-modal clinical tasks are carried out to assess model efficacy. Extensive experiments on both internal and external datasets demonstrate the superior transferability, scalability, and generalization of MG-3D, showcasing its potential in advancing feature representation for 3D medical image analysis. Code will be available: https://github.com/Xuefeng-Ni/MG-3D.

3D医学影像视觉语言预训练多粒度知识医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。