用高效扩散模型解决CT扫描视野外扩难题,速度快精度高。
Efficient Image-to-Image Schrödinger Bridge for CT Field of View Extension
- 直接学习有限视野到扩展视野的映射,跳过迭代采样。
- 真实数据上误差仅152.0 HU,比顶尖模型更准。
- 单张切片0.19秒完成,速度超传统方法700倍。
计算机断层扫描(CT)是实现内部解剖结构无创高分辨率成像的核心技术。然而,当被扫描物体超出扫描仪视野(FOV)时,投影数据会截断,导致重建不完整并在视野边界产生显著伪影。传统重建算法难以从这类数据中恢复准确解剖结构,限制了临床可靠性。深度学习方法已被用于FOV扩展,其中扩散生成模型代表了图像合成的最新进展。但传统扩散模型因迭代采样过程计算量大、推理慢。为此,我们提出基于图像到图像薛定谔桥(I²SB)扩散模型的高效CT FOV扩展框架。与从纯高斯噪声合成图像的传统扩散模型不同,I²SB学习有限视野与扩展视野图像间的直接随机映射,生成过程更可解释且可追踪,提升了重建的解剖一致性与结构保真度。I²SB在模拟噪声数据上达到49.8 HU的均方根误差(RMSE),在真实数据上为152.0 HU,优于条件去噪扩散概率模型(cDDPM)和基于块的扩散方法。此外,其单步推理可在每2D切片0.19秒内完成,相比cDDPM(135秒)提速超过700倍,优于第二快的DiffusionGAN(0.58秒)。该精度与效率的结合表明,I²SB具备实时或临床部署潜力。
原文摘要 · Abstract (English)
Computed tomography (CT) is a cornerstone imaging modality for non-invasive, high-resolution visualization of internal anatomical structures. However, when the scanned object exceeds the scanner's field of view (FOV), projection data are truncated, resulting in incomplete reconstructions and pronounced artifacts near FOV boundaries. Conventional reconstruction algorithms struggle to recover accurate anatomy from such data, limiting clinical reliability. Deep learning approaches have been explored for FOV extension, with diffusion generative models representing the latest advances in image synthesis. Yet, conventional diffusion models are computationally demanding and slow at inference due to their iterative sampling process. To address these limitations, we propose an efficient CT FOV extension framework based on the image-to-image Schrödinger Bridge (I$^2$SB) diffusion model. Unlike traditional diffusion models that synthesize images from pure Gaussian noise, I$^2$SB learns a direct stochastic mapping between paired limited-FOV and extended-FOV images. This direct correspondence yields a more interpretable and traceable generative process, enhancing anatomical consistency and structural fidelity in reconstructions. I$^2$SB achieves superior quantitative performance, with root-mean-square error (RMSE) values of 49.8 HU on simulated noisy data and 152.0 HU on real data, outperforming state-of-the-art diffusion models such as conditional denoising diffusion probabilistic models (cDDPM) and patch-based diffusion methods. Moreover, its one-step inference enables reconstruction in just 0.19 s per 2D slice, representing over a 700-fold speedup compared to cDDPM (135 s) and surpassing DiffusionGAN (0.58 s), the second fastest. This combination of accuracy and efficiency indicates that I$^2$SB has potential for real-time or clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。