arXiv:2604.18757cs.CVcs.AI2026-04中稿 · publication a MIDL…被引 1

通过视网膜与临床风险联合建模,提前8年预测阿尔茨海默病。

REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

论文配图:REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction
图 1 · 摘自论文原文
  • 将眼底图像与问卷式风险因素转化为可比语言描述,实现多模态对齐。
  • 在平均8年(1-11年)前预测新发阿尔茨海默病,显著优于现有方法。
  • 引入群体感知对比学习,强化相似患者间的跨模态关联学习。

视网膜为阿尔茨海默病(AD)和痴呆提供了独特的非侵入性观察窗口,其形态学特征可捕捉早期结构变化,而系统性和生活方式风险因素则反映了疾病易感性的长期贡献。然而,现有视网膜分析框架通常独立处理影像与风险因素,难以捕捉关键的多模态联合模式。此外,现有方法很少引入机制来组织或对齐具有相似视网膜与临床特征的患者,限制了跨模态关联的学习。为此,我们提出REVEAL(REtinal-risk Vision-Language Early Alzheimer's Learning),通过将彩色眼底照片与个体化疾病特异性风险档案对齐,实现对新发AD和痴呆的预测,平均提前8年(范围:1-11年)诊断。由于真实世界的风险因素为结构化问卷数据,我们将其转换为与预训练视觉-语言模型(VLMs)兼容的临床可解释叙述。我们进一步提出群体感知对比学习(GACL)策略,将具有相似视网膜形态学和风险因素的患者聚类为正样本对,增强多模态对齐。该统一表示学习框架显著优于最先进的视网膜影像模型搭配临床文本编码器,以及通用的VLMs,证明了联合建模视网膜生物标志物与临床风险因素的价值。通过提供一种可泛化且非侵入性的早期风险分层方法,REVEAL有望在人群中实现更早干预,提升预防性医疗水平。

原文摘要 · Abstract (English)

The retina provides a unique, noninvasive window into Alzheimer's disease (AD) and dementia, capturing early structural changes through morphometric features, while systemic and lifestyle risk factors reflect well-established contributors to disease susceptibility long before clinical symptom onset. However, current retinal analysis frameworks typically model imaging and risk factors separately, limiting their ability to capture joint multimodal patterns critical for early risk prediction. Moreover, existing methods rarely incorporate mechanisms to organize or align patients with similar retinal and clinical characteristics, constraining the learning of coherent cross-modal associations. To address these limitations, we introduce REVEAL (REtinal-risk Vision-Language Early Alzheimer's Learning), a framework that aligns color fundus photographs with individualized disease-specific risk profiles for predicting incident AD and dementia, on average 8 years before diagnosis (range: 1-11 years). Because real-world risk factors are structured questionnaire data, we translate them into clinically interpretable narratives compatible with pretrained vision-language models (VLMs). We further propose a group-aware contrastive learning (GACL) strategy that clusters patients with similar retinal morphometry and risk factors as positive pairs, strengthening multimodal alignment. This unified representation learning framework substantially outperforms state-of-the-art retinal imaging models paired with clinical text encoders, as well as general-purpose VLMs, demonstrating the value of jointly modeling retinal biomarkers and clinical risk factors. By providing a generalizable and noninvasive approach for early AD and dementia risk stratification, REVEAL has the potential to enable earlier intervention and improve preventive care at the population level.

阿尔茨海默病视网膜成像多模态学习早期预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。