arXiv:2504.08635cs.CV2025-04被引 5

用压缩空间扩散模型提升脑影像无监督学习效率与效果

Latent Diffusion Autoencoders: Toward Efficient and Meaningful Unsupervised Representation Learning in Medical Imaging

  • 在低维隐空间进行扩散建模,显著降低计算开销
  • 诊断准确率90%(AUC),年龄预测误差仅4.1年
  • 生成图像逼真且可编辑,适合长时序医学影像分析

本研究提出潜空间扩散自编码器(LDAE),一种基于扩散的新型编码器-解码器框架,用于高效且有意义的医学影像无监督学习。以阿尔茨海默病(AD)为例,利用ADNI数据库的脑部MRI数据进行验证。与传统在图像空间运行的扩散自编码器不同,LDAE在压缩的隐空间中执行扩散过程,提升了计算效率,使3D医学影像表征学习成为可能。实验验证两个假设:(i) LDAE能有效捕捉与AD和衰老相关的有意义语义表征;(ii) 实现高质量图像生成与重建,同时具备高效性。结果表明:(i) 线性探测评估显示,对AD诊断的ROC-AUC达90%,准确率84%;年龄预测的平均绝对误差(MAE)为4.1年,均方根误差(RMSE)为5.2年;(ii) 学习到的语义表征支持属性操控,生成解剖结构合理的修改;(iii) 语义插值实验显示,对缺失扫描的重建效果优异,6个月时间缺口下SSIM达0.969(MSE: 0.0019),24个月缺口下仍保持稳健性能(SSIM > 0.93,MSE < 0.004),表明其可捕捉时间演变趋势;(iv) 相比传统方法,LDAE推理吞吐量提升20倍,同时改善重建质量。这些发现表明LDAE是可扩展医学影像应用的有力框架,具有作为医疗影像分析基础模型的潜力。代码已开源。

原文摘要 · Abstract (English)

This study presents Latent Diffusion Autoencoder (LDAE), a novel encoder-decoder diffusion-based framework for efficient and meaningful unsupervised learning in medical imaging, focusing on Alzheimer disease (AD) using brain MR from the ADNI database as a case study. Unlike conventional diffusion autoencoders operating in image space, LDAE applies the diffusion process in a compressed latent representation, improving computational efficiency and making 3D medical imaging representation learning tractable. To validate the proposed approach, we explore two key hypotheses: (i) LDAE effectively captures meaningful semantic representations on 3D brain MR associated with AD and ageing, and (ii) LDAE achieves high-quality image generation and reconstruction while being computationally efficient. Experimental results support both hypotheses: (i) linear-probe evaluations demonstrate promising diagnostic performance for AD (ROC-AUC: 90%, ACC: 84%) and age prediction (MAE: 4.1 years, RMSE: 5.2 years); (ii) the learned semantic representations enable attribute manipulation, yielding anatomically plausible modifications; (iii) semantic interpolation experiments show strong reconstruction of missing scans, with SSIM of 0.969 (MSE: 0.0019) for a 6-month gap. Even for longer gaps (24 months), the model maintains robust performance (SSIM > 0.93, MSE < 0.004), indicating an ability to capture temporal progression trends; (iv) compared to conventional diffusion autoencoders, LDAE significantly increases inference throughput (20x faster) while also enhancing reconstruction quality. These findings position LDAE as a promising framework for scalable medical imaging applications, with the potential to serve as a foundation model for medical image analysis. Code available at https://github.com/GabrieleLozupone/LDAE

扩散模型医学影像无监督学习隐空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。