用零样本模型融合脑影像与常规临床数据,提前预测阿尔茨海默病风险。
Clinical Pathways Matter for Multimodal Deep Learning in Early Alzheimers Disease Detection
- 零样本框架用SigLIP提取影像与临床文本嵌入,无需任务微调
- 单次就诊时预测AUC达0.91,优于脑脊液和认知量表模型
- 支持纵向数据扩展,适合临床实用场景
识别阿尔茨海默病(AD)高危人群,尤其是无症状和轻度认知障碍(MCI)阶段,仍具挑战。尽管基于结构磁共振成像(sMRI)的深度学习方法作为非侵入性生物标志物有前景,但现有多模态模型需针对特定任务训练,且依赖临床中不常规获取的生物标志物。本文提出一种基于SigLIP的零样本多模态框架,将结构MRI嵌入与常规临床变量文本嵌入结合,用于预临床或轻度认知障碍阶段个体的早期AD风险分层。在ADNI队列416名受试者(年龄:72.73 ± 6.7岁)上评估,使用未微调的SigLIP提取MRI与临床文本嵌入,融合为多模态表示,实现4年内个体水平的风险预测。进一步比较单次与两次就诊设置下模型表现,以评估纵向信息价值与可扩展性。单次就诊下,结合MRI、MMSE、年龄和性别,AUC为0.91 ± 0.02,优于基于脑脊液Aβ42的模型(AUC 0.73 ± 0.08)和仅基于MMSE的模型(AUC 0.85 ± 0.22)。两次就诊设置下性能保持或提升,验证了该方法对纵向数据的可扩展性。结果表明,零样本多模态融合结构MRI与常规临床变量,可能为早期AD风险分层提供实用且可扩展的策略。
原文摘要 · Abstract (English)
Identifying individuals at risk of Alzheimer's disease (AD), particularly in the preclinical and early stages, remains challenging. Although deep learning approaches based on structural MRI show promise as a non-invasive biomarker, existing multimodal models require task-specific training and depend on biomarkers that are not routinely available in clinical practice. Here, we propose a zero-shot multimodal framework based on SigLIP that combines structural MRI embeddings with text embeddings of routinely collected clinical variables for early AD risk stratification in individuals at preclinical or mild cognitive impairment (MCI) stages. We evaluated the approach in 416 individuals from the ADNI cohort (age: 72.73 +- 6.7). SigLIP was used without fine-tuning to extract MRI and clinical text embeddings, which were combined into multimodal representations for individual-level AD risk prediction within 4 years. We further compared the model performance in a single-visit and two-visit settings to assess the value of longitudinal information and framework scalability. In the single-visit setting, combining MRI embeddings with MMSE, age, and sex achieved an AUC of 0.91 +- 0.02, outperforming both a CSF A\b{eta}42-based model (AUC 0.73 +- 0.08) and an MMSE-based model (AUC 0.85 +- 0.22). In the two-visit setting, performance was maintained or improved, supporting the scalability of the approach to longitudinal data. These findings suggest that zero-shot multimodal fusion of structural MRI and routinely collected clinical variables may provide a practical and scalable strategy for early AD risk stratification without task-specific retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。