arXiv:2604.22557eess.IVcs.CV2026-04中稿 · CVPR

用自然图像预训练模型提升心脏MRI加速重建的泛化能力

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?

论文配图:Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?
图 1 · 摘自论文原文
  • 将CLIP等自然图像模型作为先验嵌入重建框架
  • 跨域测试中高加速下表现优于专用模型
  • 自然图像模型迁移性更强,适合低采样场景

大规模预训练基础模型在计算机视觉中取得显著进展,但其在物理驱动的逆问题(如加速心脏MRI重建)中的潜力仍不明确。本文研究自然域基础模型能否作为有效图像先验,并与生物医学专用模型BiomedCLIP对比。提出一种可展开重建框架,将冻结的视觉编码器(如CLIP、DINOv2、BiomedCLIP)嵌入每级重构流程以引导重建。实验表明:虽然特定任务最优模型(如E2E-VarNet)在标准分布内表现更优,但在跨域场景下(心脏训练、膝关节/脑部测试),基础模型更具鲁棒性,尤其在高加速因子和低频采样受限时表现突出。自然图像预训练模型(如CLIP)学习到高度可迁移的结构表征,而领域专用预训练(BiomedCLIP)仅在病态情形下带来小幅提升。结果表明,预训练基础模型是提升加速MRI重建鲁棒性与泛化能力的有前景先验。

原文摘要 · Abstract (English)

The emergence of large-scale pretrained foundation models has transformed computer vision, enabling strong performance across diverse downstream tasks. However, their potential for physics-based inverse problems, such as accelerated cardiac MRI reconstruction, remains largely underexplored. In this work, we investigate whether natural-domain foundation models can serve as effective image priors for accelerated cardiac MRI reconstruction, and compare the performance obtained against domain-specific counterparts such as BiomedCLIP. We propose an unrolled reconstruction framework that incorporates pretrained, frozen visual encoders, such as CLIP, DINOv2, and BiomedCLIP, within each cascade to guide the reconstruction process. Through extensive experiments, we show that while task-specific state-of-the-art reconstruction models such as E2E-VarNet achieve superior performance in standard in-distribution settings, foundation-model-based approaches remain competitive. More importantly, in challenging cross-domain scenarios, where models are trained on cardiac MRI and evaluated on anatomically distinct knee and brain datasets--foundation models exhibit improved robustness, particularly under high acceleration factors and limited low-frequency sampling. We further observe that natural-image-pretrained models, such as CLIP, learn highly transferable structural representations, while domain-specific pretraining (BiomedCLIP) provides modest additional gains in more ill-posed regimes. Overall, our results suggest that pretrained foundation models offer a promising source of transferable priors, enabling improved robustness and generalization in accelerated MRI reconstruction.

MRI重建基础模型跨域泛化图像先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。