PPDM通过像素拼图机制实现3D医学图像高效生成,内存占用降低十倍。
PPDM: Pixel Puzzling Diffusion Model for Speed and Memory Efficient Volumetric Medical Image Translation

- 用可逆像素拼图操作替换空间分辨率以节省显存
- 在低计数PET去噪等任务中性能优于全量3D模型
- 适合算力受限场景下的高保真3D医学图像生成
扩散模型在医学图像到图像转换中表现出卓越保真度,但其在高分辨率3D体积数据上的应用受制于高昂的计算成本和显存需求。现有内存高效策略常牺牲全局体积一致性或精细解剖细节。本文提出像素拼图扩散模型(PPDM),一种简单高效的3D医学图像翻译框架。PPDM引入可逆像素拼图-还原算子,以通道维度替代空间分辨率,显著降低激活显存并保留全局上下文。为提升效率与稳定性,采用直接桥接扩散形式,从条件输入而非纯噪声开始,使模型聚焦于任务相关残差。此外,引入拼图梯度损失,强化空间一致性,抑制拼图带来的网格伪影。我们在多个挑战性3D医学图像翻译任务上评估了PPDM,包括低计数PET去噪、联合PET去噪与衰减校正、跨模态MRI转换。所有任务中,PPDM性能均匹配或超越完整3D扩散模型,训练显存使用最高降低一个数量级,推理速度显著提升,并优于基于潜在压缩或频域分解的现有高效方法。结果表明,PPDM为资源受限条件下高保真3D扩散医学图像转换提供了实用且可扩展的解决方案。
原文摘要 · Abstract (English)
Diffusion models have demonstrated superior fidelity for medical image-to-image translation, but their extension to high-resolution 3D volumes is severely constrained by prohibitive computational cost and GPU memory requirements. Existing memory-efficient strategies often compromise global volumetric consistency or fine anatomical detail. In this work, we propose the Pixel Puzzling Diffusion Model (PPDM), a simple and effective framework for memory- and speed-efficient 3D medical image translation. PPDM introduces a reversible pixel puzzle-unpuzzle operator that trades spatial resolution for channel dimensionality, substantially reducing activation memory while preserving global context. To further improve efficiency and stability, we adopt a direct bridge diffusion formulation that starts from the conditional input rather than pure noise, enabling the model to focus on task-relevant residuals. In addition, a puzzle-gradient loss is incorporated to enforce spatial coherence and suppress grid-like artifacts introduced by spatial rearrangement. We evaluate PPDM on multiple challenging 3D medical image translation tasks, including low-count PET denoising, joint PET denoising and attenuation correction, and cross-modal MRI translation. Across all tasks, PPDM consistently matches or outperforms full 3D diffusion models while reducing training GPU memory usage by up to an order of magnitude and significantly accelerating inference, and it outperforms existing memory-efficient diffusion approaches based on latent compression or frequency decomposition. These results demonstrate that PPDM provides a practical and scalable solution for high-fidelity 3D diffusion-based medical image translation under limited computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。