arXiv:2607.13812eess.IVcs.CV2026-07AAAI被引 4

用三平面表示降低3D医学图像生成内存占用,提升分辨率与质量。

TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

论文配图:TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model
图 1 · 摘自论文原文
  • 采用解码器仅结构学习三平面表征,减少存储开销。
  • 在脑瘤、胰腺、结肠数据集上重建与生成效果优于同类方法。
  • 适合需要高分辨3D医学图像生成的研究者使用。

我们提出TCAM-Diff,一种新型3D医学图像生成模型,可显著降低高分辨率3D数据编码与生成的内存需求。该模型采用解码器仅(decoder-only)自编码器方法,从密集体数据中学习三平面表征,并利用泛化操作防止过拟合。随后,通过三平面感知的交叉注意力扩散模型有效学习并融合这些特征。此外,扩散模型生成的特征可借助预训练解码器模块快速转换为3D体数据。我们在三个不同尺度的医学数据集上进行实验:BrainTumour(128×128×128)、Pancreas(256×256×256)和Colon(512×512×512),结果表现优异。采用均方误差(MSE)和结构相似性(SSIM)评估重建质量,使用Wasserstein生成对抗网络(W-GAN)判别器评估生成质量。与现有编码器-解码器方法相比,本方法在相似潜空间大小下取得更优的重建与生成效果。

原文摘要 · Abstract (English)

We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128 x 128 x 128, Pancreas 256 x 256 x 256, and Colon 512 x 512 x 512, demonstrate outstanding results. We utilized MSE and SSIM to assess reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons with existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces.

3D生成医学图像扩散模型三平面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。