arXiv:2507.11152eess.IVcs.AI2025-07中稿 · ed

用跨模态对齐提升稀疏视角CT重建质量

Latent Space Consistency for Sparse-View CT Reconstruction

  • 设计跨模态对比学习,统一2D X-ray与3D CT的潜在空间
  • 在LIDC-IDRI和CTSpine1K上PSNR/SSIM均超越主流模型
  • 方法可推广至文本到图像等跨模态生成任务

计算机断层扫描(CT)是临床常用成像技术,依赖密集旋转X射线阵列获取三维空间信息。然而,其面临耗时长、辐射剂量高等挑战。基于稀疏视角X射线图像的CT重建方法受到广泛关注,可降低医疗成本与风险。近年来,扩散模型尤其是潜在扩散模型(LDM)在3D CT重建中展现出潜力。但由于2D X射线与3D CT在潜在表示上的显著差异,原始LDM难以实现潜在空间对齐。为此,本文提出一致潜在空间扩散模型(CLS-DM),通过跨模态特征对比学习,从2D X射线图像中高效提取3D潜在信息,实现模态间潜在空间对齐。实验表明,CLS-DM在LIDC-IDRI与CTSpine1K数据集上,于标准体素级指标(PSNR、SSIM)上优于经典及先进生成模型。该方法不仅提升了稀疏视角重建的效率与经济性,还可推广至文本到图像合成等其他跨模态转换任务。代码已公开于https://anonymous.4open.science/r/CLS-DM-50D6/,以促进跨领域研究与应用。

原文摘要 · Abstract (English)

Computed Tomography (CT) is a widely utilized imaging modality in clinical settings. Using densely acquired rotational X-ray arrays, CT can capture 3D spatial features. However, it is confronted with challenged such as significant time consumption and high radiation exposure. CT reconstruction methods based on sparse-view X-ray images have garnered substantial attention from researchers as they present a means to mitigate costs and risks. In recent years, diffusion models, particularly the Latent Diffusion Model (LDM), have demonstrated promising potential in the domain of 3D CT reconstruction. Nonetheless, due to the substantial differences between the 2D latent representation of X-ray modalities and the 3D latent representation of CT modalities, the vanilla LDM is incapable of achieving effective alignment within the latent space. To address this issue, we propose the Consistent Latent Space Diffusion Model (CLS-DM), which incorporates cross-modal feature contrastive learning to efficiently extract latent 3D information from 2D X-ray images and achieve latent space alignment between modalities. Experimental results indicate that CLS-DM outperforms classical and state-of-the-art generative models in terms of standard voxel-level metrics (PSNR, SSIM) on the LIDC-IDRI and CTSpine1K datasets. This methodology not only aids in enhancing the effectiveness and economic viability of sparse X-ray reconstructed CT but can also be generalized to other cross-modal transformation tasks, such as text-to-image synthesis. We have made our code publicly available at https://anonymous.4open.science/r/CLS-DM-50D6/ to facilitate further research and applications in other domains.

CT重建扩散模型跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。