arXiv:2606.20112cs.CVeess.IV2026-06中稿 · ICLR被引 3

PRDiT通过分步建模实现高质量3D CT生成,突破计算瓶颈。

Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation

论文配图:Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation
图 1 · 摘自论文原文
  • 分两阶段建模:先局部去噪,再全局残差精修
  • 在LIDC-IDRI和RAD-ChestCT上3D FID等指标更优
  • 适合需要高保真医学影像生成的研究者

由于现有生成模型存在巨大的计算开销和优化困难,生成具有精细细节的高分辨率3D CT体数据仍具挑战。本文提出像素级残差扩散变压器(PRDiT),一种可扩展的生成框架,直接在体素级别合成高质量3D医学体积。PRDiT采用两阶段训练架构:1)基于MLP的盲估计器处理重叠3D块,高效分离低频结构;2)使用内存高效的注意力机制的全局残差扩散变压器,建模并精修整个体积中的高频残差。这种自粗至细的建模策略简化了优化过程,提升了训练稳定性,并有效保留微小结构,避免了自编码器瓶颈限制。在LIDC-IDRI和RAD-ChestCT数据集上的大量实验表明,PRDiT始终优于现有先进模型(如HA-GAN、3D LDM和WDM-3D),显著降低了3D FID、MMD与Wasserstein距离得分。

原文摘要 · Abstract (English)

Generating high-resolution 3D CT volumes with fine details remains challenging due to substantial computational demands and optimization difficulties inherent to existing generative models. In this paper, we propose the Pixel-Level Residual Diffusion Transformer (PRDiT), a scalable generative framework that synthesizes high-quality 3D medical volumes directly at voxel-level. PRDiT introduces a two-stage training architecture comprising 1) a local denoiser in the form of an MLP-based blind estimator operating on overlapping 3D patches to separate low-frequency structures efficiently, and 2) a global residual diffusion transformer employing memory-efficient attention to model and refine high-frequency residuals across entire volumes. This coarse-to-fine modeling strategy simplifies optimization, enhances training stability, and effectively preserves subtle structures without the limitations of an autoencoder bottleneck. Extensive experiments conducted on the LIDC-IDRI and RAD-ChestCT datasets demonstrate that PRDiT consistently outperforms state-of-the-art models, such as HA-GAN, 3D LDM and WDM-3D, achieving significantly lower 3D FID, MMD and Wasserstein distance scores.

3D生成医学影像扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。