arXiv:2502.01780stat.MLcs.LG2025-02被引 2

基于图结构改进相关性分析,提升多组学数据关联发现能力

Graph Canonical Correlation Analysis

  • 利用变量间交叉相关矩阵的图结构优化相关性计算
  • 在模拟和真实多组学数据中均显著优于传统CCA方法
  • 适合研究基因调控网络等复杂生物数据关联的科研人员

典型相关分析(CCA)是估计两组多维变量间关联的常用方法。近年来,其应用已扩展至多组学、影像组学等数据。然而,传统CCA无法有效整合交叉相关矩阵中的结构模式,可能导致估计效果不佳。为此,我们提出图式典型相关分析(gCCA),基于两组变量间交叉相关矩阵的图结构计算典型相关。开发了高效的计算算法,并通过集中不等式与鞅理论中的停止时间规则,给出了最优子集选择与典型相关估计的有限样本理论结果。大量模拟实验表明,gCCA性能优于现有方法。进一步将gCCA应用于DNA甲基化与RNA-seq多组学数据,识别出甲基化对基因表达通路的正向与负向调控关系。

原文摘要 · Abstract (English)

Canonical correlation analysis (CCA) is a widely used technique for estimating associations between two sets of multi-dimensional variables. Recent advancements in CCA methods have expanded their application to decipher the interactions of multiomics datasets, imaging-omics datasets, and more. However, conventional CCA methods are limited in their ability to incorporate structured patterns in the cross-correlation matrix, potentially leading to suboptimal estimations. To address this limitation, we propose the graph Canonical Correlation Analysis (gCCA) approach, which calculates canonical correlations based on the graph structure of the cross-correlation matrix between the two sets of variables. We develop computationally efficient algorithms for gCCA, and provide theoretical results for finite sample analysis of best subset selection and canonical correlation estimation by introducing concentration inequalities and stopping time rule based on martingale theories. Extensive simulations demonstrate that gCCA outperforms competing CCA methods. Additionally, we apply gCCA to a multiomics dataset of DNA methylation and RNA-seq transcriptomics, identifying both positively and negatively regulated gene expression pathways by DNA methylation pathways.

多组学分析典型相关图模型生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。