arXiv:2509.06617eess.IVcs.CV2025-09被引 11

将DINOv2适配多模态医学影像,提升脑瘤分类准确率。

MM-DINOv2: Adapting Foundation Models for Multi-Modal Medical Image Analysis

  • 引入多模态图像块嵌入,让模型处理多序列MRI
  • 用全模态掩码应对缺失数据,提升鲁棒性
  • 半监督学习利用大量无标签数据,适合临床实际

如DINOv2等视觉基础模型在医学影像中展现出巨大潜力,但其设计主要针对单模态图像,难以胜任神经科和肿瘤科常见的多模态任务。虽有监督模型表现良好,却无法利用无标签数据,且对缺失模态敏感。为此,我们提出MM-DINOv2,一种高效框架,将预训练的DINOv2适配多模态医学影像。该方法引入多模态图像块嵌入,使模型能有效处理多序列数据;采用全模态掩码策略,增强跨模态关系学习能力;结合半监督学习,充分利用大规模无标签数据。在多序列脑MRI胶质瘤亚型分类任务中,外部测试集MCC达0.6,较现有最优监督方法提升11.1%。本工作为多模态医学影像提供可扩展、稳健的解决方案,兼顾自然图像预训练模型优势与真实临床挑战。

原文摘要 · Abstract (English)

Vision foundation models like DINOv2 demonstrate remarkable potential in medical imaging despite their origin in natural image domains. However, their design inherently works best for uni-modal image analysis, limiting their effectiveness for multi-modal imaging tasks that are common in many medical fields, such as neurology and oncology. While supervised models perform well in this setting, they fail to leverage unlabeled datasets and struggle with missing modalities, a frequent challenge in clinical settings. To bridge these gaps, we introduce MM-DINOv2, a novel and efficient framework that adapts the pre-trained vision foundation model DINOv2 for multi-modal medical imaging. Our approach incorporates multi-modal patch embeddings, enabling vision foundation models to effectively process multi-modal imaging data. To address missing modalities, we employ full-modality masking, which encourages the model to learn robust cross-modality relationships. Furthermore, we leverage semi-supervised learning to harness large unlabeled datasets, enhancing both the accuracy and reliability of medical predictions. Applied to glioma subtype classification from multi-sequence brain MRI, our method achieves a Matthews Correlation Coefficient (MCC) of 0.6 on an external test set, surpassing state-of-the-art supervised approaches by +11.1%. Our work establishes a scalable and robust solution for multi-modal medical imaging tasks, leveraging powerful vision foundation models pre-trained on natural images while addressing real-world clinical challenges such as missing data and limited annotations.

多模态医学影像DINOv2半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。