arXiv:2512.23441cs.LGcs.CV2025-12被引 1

用随机过程建模医学影像时间动态,提升疾病进展预测能力

Stochastic Siamese MAE Pretraining for Longitudinal Medical Images

  • 设计双分支MAE框架,通过时间差条件建模时序不确定性
  • 在OCT与MRI数据上,对晚期黄斑变性等疾病预测效果更优
  • 适合需要捕捉非确定性病程的医学影像分析任务

时间感知的图像表征对于捕捉纵向医学数据中疾病进展至关重要。然而,当前最先进的自监督学习方法如掩码自编码器(MAE),尽管具备强大的表征学习能力,却缺乏时间感知。本文提出STAMP(随机时间自编码器与掩码预训练),一种基于双分支MAE的框架,通过条件化两个输入体积的时间差,以随机过程编码时间信息。与传统确定性双分支方法不同,后者虽比较不同时期扫描但无法考虑疾病演化的内在不确定性,STAMP将MAE重建损失重构为条件变分推断目标,从而随机地学习时间动态。我们在两个OCT和一个MRI数据集上评估了STAMP,这些数据集每位患者有多次随访。使用STAMP预训练的ViT模型,在不同晚期年龄相关性黄斑变性和阿尔茨海默病进展预测任务中,优于现有时间感知的MAE方法及基础模型,证明其在学习疾病非确定性时间动态方面的优势。

原文摘要 · Abstract (English)

Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets. However, recent state-of-the-art self-supervised learning approaches like Masked Autoencoding (MAE), despite their strong representation learning capabilities, lack temporal awareness. In this paper, we propose STAMP (Stochastic Temporal Autoencoder with Masked Pretraining), a Siamese MAE framework that encodes temporal information through a stochastic process by conditioning on the time difference between the 2 input volumes. Unlike deterministic Siamese approaches, which compare scans from different time points but fail to account for the inherent uncertainty in disease evolution, STAMP learns temporal dynamics stochastically by reframing the MAE reconstruction loss as a conditional variational inference objective. We evaluated STAMP on two OCT and one MRI datasets with multiple visits per patient. STAMP pretrained ViT models outperformed both existing temporal MAE methods and foundation models on different late stage Age-Related Macular Degeneration and Alzheimer's Disease progression prediction which require models to learn the underlying non-deterministic temporal dynamics of the diseases.

自监督学习医学影像时间建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。