用临床文本控制阿尔茨海默病影像生成,精准模拟个体进展过程。
ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression

- 将随访时间与多维度临床信息融合为自然语言提示,实现时间精细化控制。
- 在3321张纵向MRI上达成SSIM 0.8739、PSNR 29.32 dB,性能优于基线模型。
- 适合需要个性化疾病进展模拟的临床研究与辅助诊断场景。
阿尔茨海默病(AD)在个体间进展异质性强,推动了基于受试者特异性合成随访磁共振成像(MRI)以支持病情评估的需求。尽管扩散变换器(DiT)作为新兴的基于变压器的扩散模型,为图像合成提供了可扩展的骨干架构,但对具有临床可解释性的时间间隔和参与者元数据控制的纵向AD MRI生成仍研究不足。本文提出ADP-DiT,一种区间感知、临床文本条件化的扩散变换器,用于纵向AD MRI合成。ADP-DiT将随访间隔及多领域人口统计学、诊断状态(CN/MCI/AD)和神经心理学信息编码为自然语言提示,实现超越粗略诊断阶段的时间特定控制。为有效注入这种条件,采用双文本编码器——OpenCLIP用于视觉-语言对齐,T5用于更丰富的临床语言理解。其嵌入通过交叉注意力融合至DiT,并使用自适应层归一化进行全局调制。为进一步提升解剖保真度,对图像标记应用旋转位置编码,并在预训练的SDXL-VAE潜在空间中进行扩散,以实现高效高分辨率重建。在来自712名参与者(共259,038个图像切片)的3,321个纵向3T T1加权扫描上,ADP-DiT达到SSIM 0.8739和PSNR 29.32 dB,较DiT基线分别提升+0.1087和+6.08 dB,且能捕捉如脑室扩大和海马萎缩等进展相关变化。结果表明,将全面的个体化临床条件与先进架构结合,可显著提升纵向AD MRI合成效果。
原文摘要 · Abstract (English)
Alzheimer's disease (AD) progresses heterogeneously across individuals, motivating subject-specific synthesis of follow-up magnetic resonance imaging (MRI) to support progression assessment. While Diffusion Transformers (DiT), an emerging transformer-based diffusion model, offer a scalable backbone for image synthesis, longitudinal AD MRI generation with clinically interpretable control over follow-up time and participant metadata remains underexplored. We present ADP-DiT, an interval-aware, clinically text-conditioned diffusion transformer for longitudinal AD MRI synthesis. ADP-DiT encodes follow-up interval together with multi-domain demographic, diagnostic (CN/MCI/AD), and neuropsychological information as a natural-language prompt, enabling time-specific control beyond coarse diagnostic stages. To inject this conditioning effectively, we use dual text encoders-OpenCLIP for vision-language alignment and T5 for richer clinical-language understanding. Their embeddings are fused into DiT through cross-attention for fine-grained guidance and adaptive layer normalization for global modulation. We further enhance anatomical fidelity by applying rotary positional embeddings to image tokens and performing diffusion in a pre-trained SDXL-VAE latent space to enable efficient high-resolution reconstruction. On 3,321 longitudinal 3T T1-weighted scans from 712 participants (259,038 image slices), ADP-DiT achieves SSIM 0.8739 and PSNR 29.32 dB, improving over a DiT baseline by +0.1087 SSIM and +6.08 dB PSNR while capturing progression-related changes such as ventricular enlargement and shrinking hippocampus. These results suggest that integrating comprehensive, subject-specific clinical conditions with architectures can improve longitudinal AD MRI synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。