arXiv:2507.16579eess.IVcs.AI2025-07

用分层掩码扩散模型,高效合成高质量医学影像。

Pyramid Hierarchical Masked Diffusion Model for Imaging Synthesis

  • 分层多尺度掩码加速训练,控制不同分辨率的生成细节。
  • 在两个数据集上PSNR和SSIM均优于现有方法。
  • 适合需要跨模态医学图像补全的研究者使用。

医学图像合成在临床工作中至关重要,可解决因扫描时间长、图像伪影、患者运动或对对比剂不耐受导致的模态缺失问题。本文提出一种新型图像合成网络——金字塔分层掩码扩散模型(PHMDiff),采用多尺度分层策略,在不同分辨率和层级上实现更精细的控制。该模型利用随机多尺度高比例掩码以加速扩散模型训练,并平衡细节保真度与整体结构。通过引入基于Transformer的扩散过程,结合跨粒度正则化,建模各粒度潜在空间间的互信息一致性,从而提升像素级感知准确性。在两个挑战性数据集上的综合实验表明,PHMDiff在峰值信噪比(PSNR)和结构相似性指数(SSIM)方面均取得更优性能,展现出生成高质量且结构完整的合成图像能力。消融实验进一步验证了各组件的有效性。PHMDiff作为跨模态及模态内多尺度图像合成框架,显著优于其他方法。源代码已公开于https://github.com/xiaojiao929/PHMDiff。

原文摘要 · Abstract (English)

Medical image synthesis plays a crucial role in clinical workflows, addressing the common issue of missing imaging modalities due to factors such as extended scan times, scan corruption, artifacts, patient motion, and intolerance to contrast agents. The paper presents a novel image synthesis network, the Pyramid Hierarchical Masked Diffusion Model (PHMDiff), which employs a multi-scale hierarchical approach for more detailed control over synthesizing high-quality images across different resolutions and layers. Specifically, this model utilizes randomly multi-scale high-proportion masks to speed up diffusion model training, and balances detail fidelity and overall structure. The integration of a Transformer-based Diffusion model process incorporates cross-granularity regularization, modeling the mutual information consistency across each granularity's latent spaces, thereby enhancing pixel-level perceptual accuracy. Comprehensive experiments on two challenging datasets demonstrate that PHMDiff achieves superior performance in both the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), highlighting its capability to produce high-quality synthesized images with excellent structural integrity. Ablation studies further confirm the contributions of each component. Furthermore, the PHMDiff model, a multi-scale image synthesis framework across and within medical imaging modalities, shows significant advantages over other methods. The source code is available at https://github.com/xiaojiao929/PHMDiff

医学图像扩散模型图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。