arXiv:2603.12351stat.MLcs.LG2026-03

ProJIVE统一建模多组学数据的共性和特异性变化,提升解析精度。

Probabilistic Joint and Individual Variation Explained (ProJIVE) for Data Integration

  • 基于概率框架的EM算法,联合估计多数据集的共享与独特变异
  • 在阿尔茨海默病数据中,联合变异得分与昂贵生物标志物高度相关
  • 适用于基因组、神经影像等多模态数据整合,生物学意义强

在现代科学研究中,对同一组受试者收集多种类型数据(如基因组学、代谢组学、神经影像)十分常见。联合与个体变异解释(JIVE)旨在对多个特征集间的共同变化进行低秩近似,并分离出每组特征独有的变异。本文提出一种基于期望-最大化(EM)算法的概率模型,将概率主成分分析扩展至多数据集场景。该最大似然方法可同时估计联合与个体成分,相比其他方法更精确。我们在阿尔茨海默病的脑形态测量和认知指标数据上应用ProJIVE,结果揭示了具有生物学意义的变异轨迹;联合形态与认知的受试者得分与更昂贵的现有生物标志物显著相关。本研究使用的数据来自阿尔茨海默病神经影像计划(ADNI)数据库。分析代码已公开于GitHub。

原文摘要 · Abstract (English)

Collecting multiple types of data on the same set of subjects is common in modern scientific applications including, genomics, metabolomics, and neuroimaging. Joint and Individual Variance Explained (JIVE) seeks a low-rank approximation of the joint variation between two or more sets of features captured on common subjects and isolates this variation from that unique to eachset of features. We develop an expectation-maximization (EM) algorithm to estimate a probabilistic model for the JIVE framework. The model extends probabilistic principal components analysis to multiple data sets. Our maximum likelihood approach simultaneously estimates joint and individual components, which can lead to greater accuracy compared to other methods. We apply ProJIVE to measures of brain morphometry and cognition in Alzheimer's disease. ProJIVE learns biologically meaningful courses of variation, and the joint morphometry and cognition subject scores are strongly related to more expensive existing biomarkers. Data used in preparation of this article were obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. Code to reproduce the analysis is available on our GitHub page.

数据融合多组学概率模型阿尔茨海默病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。