arXiv:2507.11569eess.IVcs.AI2025-07被引 7

评估视觉大模型在乳腺MRI配准中的表现,发现其全局对齐强但细节差。

Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?

  • 用五个预训练编码器测试乳腺MRI配准任务,覆盖多时点、多模态、病灶差异
  • SAM模型在整体配准上优于传统方法,尤其在跨域场景下,但难以捕捉纤维腺体组织细节
  • 医学数据微调未提升性能,甚至降低效果,提示领域训练需更精细策略

视觉基础模型在零样本图像配准中展现出潜力,但其在刚性或结构简单器官(如脑部或腹部)上的表现已广泛研究。然而,这些模型能否应对更具挑战性的可变形解剖结构仍不明确。乳腺MRI配准尤为困难,因患者间解剖差异大、体位导致形变、且纤维腺体组织细密复杂,精确对齐至关重要。本研究系统评估了五种基于基础模型的配准算法(DINO-v2、SAM、MedSAM、SSLSAM、MedCLIP),在四个关键任务中测试:不同年份与时间点、序列、模态及疾病状态(有无病灶)。结果表明,如SAM等模型在整体乳腺对齐上优于传统基线,尤其在大域偏移下表现更优;但对纤维腺体组织的精细结构仍难准确捕捉。值得注意的是,对MedSAM和SSLSAM进行医学或乳腺特异性数据的额外预训练或微调,并未提升性能,部分情况下反而下降。这表明领域训练对配准的影响机制尚需深入理解,未来需探索能同时提升全局对齐与细结构精度的针对性策略。代码已开源至GitHub。

原文摘要 · Abstract (English)

Foundation models, pre-trained on large image datasets and capable of capturing rich feature representations, have recently shown potential for zero-shot image registration. However, their performance has mostly been tested in the context of rigid or less complex structures, such as the brain or abdominal organs, and it remains unclear whether these models can handle more challenging, deformable anatomy. Breast MRI registration is particularly difficult due to significant anatomical variation between patients, deformation caused by patient positioning, and the presence of thin and complex internal structure of fibroglandular tissue, where accurate alignment is crucial. Whether foundation model-based registration algorithms can address this level of complexity remains an open question. In this study, we provide a comprehensive evaluation of foundation model-based registration algorithms for breast MRI. We assess five pre-trained encoders, including DINO-v2, SAM, MedSAM, SSLSAM, and MedCLIP, across four key breast registration tasks that capture variations in different years and dates, sequences, modalities, and patient disease status (lesion versus no lesion). Our results show that foundation model-based algorithms such as SAM outperform traditional registration baselines for overall breast alignment, especially under large domain shifts, but struggle with capturing fine details of fibroglandular tissue. Interestingly, additional pre-training or fine-tuning on medical or breast-specific images in MedSAM and SSLSAM, does not improve registration performance and may even decrease it in some cases. Further work is needed to understand how domain-specific training influences registration and to explore targeted strategies that improve both global alignment and fine structure accuracy. We also publicly release our code at \href{https://github.com/mazurowski-lab/Foundation-based-reg}{Github}.

医学影像图像配准大模型应用乳腺MRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。