arXiv:2604.06849cs.CV2026-04

用视觉语言模型指导快速磁共振成像,实现个性化诊断优化

Vision-Language Model-Guided Deep Unrolling Enables Personalized, Fast MRI

  • 用视觉语言模型引导深度展开重建,动态生成患者特异性扫描轨迹
  • 在多种解剖结构和加速因子下,图像质量显著优于传统方法
  • 适合临床需精准检测异常的场景,如肿瘤定位与细粒度诊断

磁共振成像(MRI)是医学诊疗的核心技术,但存在采集时间长的问题。传统加速MRI方法优化通用图像质量,缺乏针对特定临床任务的适应性。为此,我们提出PASS(个性化、异常感知采样与重建)框架,利用视觉语言模型(VLM)引导深度展开网络,实现面向任务的快速成像。PASS通过三个核心创新实现个性化:(1) 基于物理模型的深度展开重建网络;(2) 生成患者特异性$k$-空间轨迹的采样模块;(3) 从预训练VLM提取的异常感知先验,指导采样与重建聚焦临床相关区域。通过融合VLM的高层临床推理与可解释的物理感知网络,PASS在多样解剖结构、对比度、异常类型及加速因子下均实现更优图像质量,直接提升下游诊断性能,包括细粒度异常检测、定位与诊断能力。

原文摘要 · Abstract (English)

Magnetic Resonance Imaging (MRI) is a cornerstone in medicine and healthcare but suffers from long acquisition times. Traditional accelerated MRI methods optimize for generic image quality, lacking adaptability for specific clinical tasks. To address this, we introduce PASS (Personalized, Anomaly-aware Sampling and reconStruction), an intelligent MRI framework that leverages a Vision-Language Model (VLM) to guide a deep unrolling network for task-oriented, fast imaging. PASS dynamically personalizes the imaging pipeline through three core contributions: (1) a deep unrolled reconstruction network derived from a physics-based MRI model; (2) a sampling module that generates patient-specific $k$-space trajectories; and (3) an anomaly-aware prior, extracted from a pretrained VLM, which steers both sampling and reconstruction toward clinically relevant regions. By integrating the high-level clinical reasoning of a VLM with an interpretable, physics-aware network, PASS achieves superior image quality across diverse anatomies, contrasts, anomalies, and acceleration factors. This enhancement directly translates to improvements in downstream diagnostic tasks, including fine-grained anomaly detection, localization, and diagnosis.

磁共振成像视觉语言模型个性化医疗医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。