arXiv:2412.09331eess.IVcs.CV2024-12被引 40

用状态空间模型实现医学影像高效高保真重建

Physics-Driven Autoregressive State Space Models for Medical Image Reconstruction

  • 基于物理驱动的自回归状态空间模型,分层递进重建图像
  • 在加速MRI和稀疏视角CT上超越现有最佳方法
  • 适合需要高精度重建的医学影像研究者

从欠采样数据中进行医学图像重建是一个病态逆问题,需从不完整测量中准确恢复解剖结构。物理驱动(PD)网络模型通过整合数据一致性机制与学习先验,在性能上优于纯数据驱动方法。然而,重建质量仍依赖于网络区分伪影与真实解剖信号的能力——二者均具有复杂的多尺度上下文结构。卷积神经网络(CNN)捕捉局部相关性但常难以建模非局部依赖;而变压器虽试图缓解此问题,但实际应用中为降低计算成本需权衡局部与非局部敏感性,有时性能仅相当于CNN。为此,我们提出MambaRoll,一种新型物理驱动的自回归状态空间模型(SSM),用于高保真、高效的图像重建。该模型采用展开式架构,每级级联自回归地预测细尺度特征图,条件于粗尺度表示,实现一致的多尺度上下文传播。每一阶段由一系列特定尺度的PD-SSM模块构成,捕捉空间依赖并通过对残差校正强制数据一致性。为进一步提升尺度感知学习,我们引入深度多尺度解码(DMSD)损失,在中间空间尺度提供监督,与自回归设计对齐。在加速MRI和稀疏视角CT重建任务中的实验表明,MambaRoll始终优于当前最先进的基于CNN、Transformer和SSM的方法。

原文摘要 · Abstract (English)

Medical image reconstruction from undersampled acquisitions is an ill-posed inverse problem requiring accurate recovery of anatomical structures from incomplete measurements. Physics-driven (PD) network models have gained prominence for this task by integrating data-consistency mechanisms with learned priors, enabling improved performance over purely data-driven approaches. However, reconstruction quality still hinges on the network's ability to disentangle artifacts from true anatomical signals-both of which exhibit complex, multi-scale contextual structure. Convolutional neural networks (CNNs) capture local correlations but often struggle with non-local dependencies. While transformers aim to alleviate this limitation, practical implementations involve design compromises to reduce computational cost by balancing local and non-local sensitivity, occasionally resulting in performance comparable to CNNs. To address these challenges, we propose MambaRoll, a novel physics-driven autoregressive state space model (SSM) for high-fidelity and efficient image reconstruction. MambaRoll employs an unrolled architecture where each cascade autoregressively predicts finer-scale feature maps conditioned on coarser-scale representations, enabling consistent multi-scale context propagation. Each stage is built on a hierarchy of scale-specific PD-SSM modules that capture spatial dependencies while enforcing data consistency through residual correction. To further improve scale-aware learning, we introduce a Deep Multi-Scale Decoding (DMSD) loss, which provides supervision at intermediate spatial scales in alignment with the autoregressive design. Demonstrations on accelerated MRI and sparse-view CT reconstructions show that MambaRoll consistently outperforms state-of-the-art CNN-, transformer-, and SSM-based methods.

医学影像状态空间模型图像重建物理驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。