用多结构差异建模+时间感知注意力,提升青光眼早期进展预测准确率。
DiffSight-Former: Modeling Structural Differences and Temporal Dynamics for Glaucoma Progression Prediction

- 基于眼底图像基础模型提取动态特征,融合结构与血流变化差异
- 在两个数据集上达到91.54%和87.48%的准确率,显著优于现有方法
- 适合临床长期随访监测,尤其适用于早期风险预警
青光眼是全球导致不可逆失明的主要原因,早期通过眼底图像检测对疾病管理至关重要。尽管深度学习在眼底图像分析中表现优异,但多数方法仅依赖单时点图像,难以捕捉疾病进展相关的结构与血管演变。临床随访获取的序列眼底图像包含重要时间信息,但现有模型常难以识别细微早期信号,且多依赖固定长度输入或已确诊图像的诊断线索,限制其早期预测能力。为此,我们提出DiffSight-Former框架,用于从序列眼底图像中预测青光眼进展。该框架采用基于眼底专用基础模型的时变特征提取模块,获得鲁棒的解剖学表征;引入多结构差异建模模块,量化视盘/杯区与视网膜血管的变化;结合时间间隔嵌入,经由时间感知Transformer建模疾病进展并估计未来发病概率。在两个纵向数据集SIGF(405个序列)和GRAPE(263个序列)上进行实验,DiffSight-Former在SIGF上达到91.54% AUC和92.16%敏感度,在GRAPE上跨三种视野进展标准平均准确率达87.48%。相比现有方法,其在不同时间设置下均表现出强性能与鲁棒性,展现出在长期青光眼监测与早期风险预测中的潜力。
原文摘要 · Abstract (English)
Glaucoma is a leading cause of irreversible blindness worldwide, and early detection from fundus images is critical for effective disease management. While deep learning has achieved promising performance in fundus image analysis, most existing methods rely on single time-point images and fail to capture longitudinal structural and vascular changes associated with disease progression. Sequential fundus images acquired during clinical follow-up provide valuable temporal information; however, current sequential models often struggle to detect subtle early progression signals and commonly depend on fixed-length inputs or diagnostic cues from already glaucomatous images, limiting their clinical utility for early prediction. To address these limitations, we propose DiffSight-Former, a framework for glaucoma progression prediction from sequential fundus images. It incorporates a time-variant feature extraction module based on a fundus-specific foundation model to obtain robust anatomical representations. A multi-structure difference modeling module is introduced to quantify progression-related changes in the optic disc/cup region and retinal vasculature. These representations are integrated with temporal interval embeddings and processed by a time-aware Transformer to model disease progression and estimate the probability of future glaucoma onset. Experiments were conducted on two longitudinal datasets, SIGF (405 sequences) and GRAPE (263 sequences). On SIGF, DiffSight-Former achieved an AUC of 91.54% and a sensitivity of 92.16% for progression prediction. On GRAPE, it achieved an average accuracy of 87.48% across three clinical visual-field progression criteria. Compared with existing approaches, DiffSight-Former demonstrates strong performance and robustness across different temporal settings, highlighting its potential for longitudinal glaucoma monitoring and early risk prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。