从多研究基因数据中分离共用与特异因子,揭示疾病相关生物通路。
Nonlinear multi-study sparse factor analysis
- 基于稀疏变分自编码器,建模共享与研究特异因子的非线性关系。
- 在血小板基因表达数据中识别出与疾病相关的共用及特异性基因簇。
- 方法可区分因子的生物学意义,适合多组学数据整合分析。
高维数据常蕴含低维潜在因子。对于来自多个研究的高维数据,目标之一是识别所有研究共有的潜在因子以及研究特异因子。以不同疾病组患者的血小板基因表达数据为例,因子对应于共表达的基因簇;我们预期某些基因簇(或生物通路)在所有疾病中均活跃,而另一些仅在特定疾病中活跃。为此,我们提出一种非线性多研究稀疏因子模型,支持共享与研究特异因子。为拟合该模型,我们设计了多研究稀疏变分自编码器。模型具有稀疏性,即每个观测特征(数据的每个维度)仅依赖少数潜在因子;在基因组学中,意味着每个基因仅参与少数生物过程。我们证明了潜在因子分布可辨识(至元素级变换),且典型共享与研究特异的支持结构可辨识。实证表明,该方法能有效恢复血小板基因表达数据中的有意义因子。
原文摘要 · Abstract (English)
High-dimensional data often exhibit variation that can be captured by lower-dimensional factors. For high-dimensional data from multiple studies, one goal is to understand which underlying factors are common to all studies, and which factors are study-specific. As a particular example, we consider platelet gene expression data from patients in different disease groups. In this data, factors correspond to clusters of genes which are co-expressed; we may expect some clusters (or biological pathways) to be active for all diseases, while some clusters are only active for a specific disease. To learn these factors, we consider a nonlinear multi-study sparse factor model, which allows for both shared and study-specific factors. To fit this model, we propose a multi-study sparse variational autoencoder. The underlying model is sparse in that each observed feature (i.e. each dimension of the data) depends on a small subset of the latent factors. In the genomics example, this means each gene is active in only a few biological processes. We prove that the latent factor distributions are identifiable (up to element-wise transformations), and that the canonical shared and study-specific support structure is identifiable. Empirically, we demonstrate our method recovers meaningful factors in the platelet gene expression data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。