arXiv:2511.19641cs.CVcs.AI2025-11被引 3

用视觉语言模型提升快速MRI重建质量,让图像更符合解剖结构和感知常识。

On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction

  • 利用预训练视觉语言模型提取图像与辅助信息的高层语义特征。
  • 相比传统正则化,重建图像在细节保留和感知质量上显著提升,LPIPS更低,Tenengrad更高。
  • 支持多模态语义先验,适合医学影像中需高精度结构还原的场景。

目的:探究视觉语言基础模型是否可通过提供超越传统先验的高层次上下文信息,提升欠采样MRI重建效果。方法:提出一种语义分布引导的重建框架,利用预训练视觉语言基础模型将重建图像及辅助信息编码为高层语义特征,并通过对比目标使重建表示与目标语义分布对齐,确保与高层次感知线索一致。该目标可与多种深度学习重建方法结合,灵活融入多模态来源的语义先验。为验证语义先验有效性,分别测试了仅图像或图像-语言辅助信息引导的重建结果。结果:在膝关节和脑部数据集上的实验表明,仅图像的语义先验能更好保留精细解剖结构,感知质量更优,表现为更低的LPIPS值、更高的Tenengrad分数及读者评估得分,优于传统正则化。图像-语言信息进一步扩展语义空间,实现对重建属性的高层控制。所有评估中,对比目标均有效引导重建特征向期望语义分布靠拢,同时保持数据保真度,验证了优化框架的有效性。结论:研究显示,视觉语言基础模型可通过语义空间优化改善欠采样MRI重建效果。

原文摘要 · Abstract (English)

Purpose: To investigate whether a vision-language foundation model can enhance undersampled MRI reconstruction by providing high-level contextual information beyond conventional priors. Methods: We proposed a semantic distribution-guided reconstruction framework that uses a pre-trained vision-language foundation model to encode both the reconstructed image and auxiliary information into high-level semantic features. A contrastive objective aligns the reconstructed representation with the target semantic distribution, ensuring consistency with high-level perceptual cues. The proposed objective works with various deep learning-based reconstruction methods and can flexibly incorporate semantic priors from multimodal sources. To test the effectiveness of these semantic priors, we evaluated reconstruction results guided by priors derived from either image-only or image-language auxiliary information. Results: Experiments on knee and brain datasets demonstrate that semantic priors from images preserve fine anatomical structures and achieve superior perceptual quality, as reflected in lower LPIPS values, higher Tenengrad scores, and improved scores in the reader study, compared with conventional regularization. The image-language information further expands the semantic distribution and enables high-level control over reconstruction attributes. Across all evaluations, the contrastive objective consistently guided the reconstructed features toward the desired semantic distributions while maintaining data fidelity, demonstrating the effectiveness of the proposed optimization framework. Conclusion: The study highlights that vision-language foundation models can improve undersampled MRI reconstruction through semantic-space optimization.

MRI重建视觉语言模型语义引导医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。