用分层扩散变换器预测脑部MRI长期变化,兼顾整体结构与局部病程细节。
ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction

- 分粗细两条路径:粗路径建模整体结构,细路径在局部精细修正。
- 直接在体素空间估计速度场,端到端生成未来脑扫描,无需自编码器。
- 适用于阿尔茨海默病等神经退行性疾病进展预测,适合医学影像研究者。
预测未来脑部结构磁共振成像具有挑战性,因纵向变化通常细微且局限于特定解剖区域,而大多数个体脑结构在时间上保持稳定。有效模型需在保持全局脑结构一致性的同时,对细微疾病进展敏感。现有基于潜在空间的方法虽提升计算效率,但在压缩-重建过程中损失信息;而直接体素空间方法常使用统一预测路径,导致稳定的脑结构掩盖细微局部变化。为此,本文提出ProgFormer,一种用于纵向脑部MRI预测的分层体素空间扩散变换器。该模型通过粗路径从3D块标记中进行主要体积预测,建模整体脑结构与纵向上下文;细路径则利用粗表示作为时空基础,在各块内进行体素级精细化调整。两条路径联合通过条件流匹配直接在体素空间估计速度场,实现无需独立训练图像自编码器的端到端预测。最终,通过一系列欧拉步积分速度场,从高斯噪声生成未来扫描。在ADNI、AIBL和OASIS三个常用基准数据集上,无论成对或轨迹设置下,均优于多种先进方法。
原文摘要 · Abstract (English)
Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, while most subject-specific brain structure remains stable over time. An effective model should therefore preserve global brain structural consistency while remaining sensitive to fine-grained disease progression. Existing latent-space-based methods improve computational efficiency, but suffer from information loss during their compression-reconstruction procedure. In contrast, direct voxel-space methods avoid latent reconstruction but commonly use a unified prediction pathway to model brain structure and progression-related changes. Subtle local changes may therefore be overshadowed by the dominant stable brain structure. To address these challenges, we propose ProgFormer, a hierarchical voxel-space Diffusion Transformer for longitudinal brain MRI prediction. ProgFormer uses a coarse pathway to perform the primary volumetric prediction from 3D patch tokens. This pathway models overall brain structure and longitudinal context. The fine pathway then uses the coarse representations as spatio-temporal grounding for voxel-level refinement within individual patches. The two pathways jointly estimate a velocity field directly in voxel space through conditional flow matching, enabling end-to-end prediction without a separately learned image autoencoder. The predicted future scan is then generated from Gaussian noise by integrating the estimated velocity field over a sequence of Euler steps. Extensive experimental results on three widely used benchmarks, ADNI, AIBL, and OASIS, under both pairwise and trajectory settings demonstrate favourable performance compared against several state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。