arXiv:2512.17276cs.LG2025-12

用半监督方法从少量标注数据中挖掘阿尔茨海默病脑网络,诊断准确率接近完美。

Alzheimer's Disease Brain Network Mining

  • 融合深度学习与最优传输理论,通过图传播实现小样本标签扩展。
  • 在近5000人数据上仅用不足1/3标注,诊断准确率近乎完美。
  • 适合临床部署,对标注稀缺场景鲁棒,理论可保证误差边界。

阿尔茨海默病(AD)的机器学习诊断面临根本挑战:临床评估昂贵且侵入性强,导致神经影像数据集仅有部分具备真实标签。我们提出多视图自适应传输聚类框架MATCH-AD,结合深度表征学习、基于图的标签传播与最优传输理论,解决这一限制。该框架利用神经影像数据的流形结构,将有限标注样本的诊断信息传播至更大未标注群体,并使用Wasserstein距离量化不同认知状态间的疾病进展。在来自国家阿尔茨海默病协调中心的近5000名受试者数据上进行评估,涵盖数百个脑区的结构磁共振成像、脑脊液生物标志物及临床变量。尽管真实标签不足全体的三分之一,MATCH-AD仍实现近乎完美的诊断准确率。其性能显著优于所有基线方法,κ值显示几乎完美一致,而最佳基线仅为弱一致性,诊断可靠性实现质变。即使在极端标签稀缺下,性能仍具临床实用性,并提供了理论上的收敛性保证,包括标签传播误差与传输稳定性的严格上界。结果表明,严谨的半监督学习可释放全球大量部分标注神经影像数据的诊断潜力,大幅降低标注负担,同时保持适用于临床部署的高精度。

原文摘要 · Abstract (English)

Machine learning approaches for Alzheimer's disease (AD) diagnosis face a fundamental challenges. Clinical assessments are expensive and invasive, leaving ground truth labels available for only a fraction of neuroimaging datasets. We introduce Multi view Adaptive Transport Clustering for Heterogeneous Alzheimer's Disease (MATCH-AD), a semi supervised framework that integrates deep representation learning, graph-based label propagation, and optimal transport theory to address this limitation. The framework leverages manifold structure in neuroimaging data to propagate diagnostic information from limited labeled samples to larger unlabeled populations, while using Wasserstein distances to quantify disease progression between cognitive states. Evaluated on nearly five thousand subjects from the National Alzheimer's Coordinating Center, encompassing structural MRI measurements from hundreds of brain regions, cerebrospinal fluid biomarkers, and clinical variables MATCHAD achieves near-perfect diagnostic accuracy despite ground truth labels for less than one-third of subjects. The framework substantially outperforms all baseline methods, achieving kappa indicating almost perfect agreement compared to weak agreement for the best baseline, a qualitative transformation in diagnostic reliability. Performance remains clinically useful even under severe label scarcity, and we provide theoretical convergence guarantees with proven bounds on label propagation error and transport stability. These results demonstrate that principled semi-supervised learning can unlock the diagnostic potential of the vast repositories of partially annotated neuroimaging data accumulating worldwide, substantially reducing annotation burden while maintaining accuracy suitable for clinical deployment.

阿尔茨海默病半监督学习脑网络分析医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。