提出可处理百万变量的多组学网络模型,高效捕捉生物调控关系。
Learning Massive-scale Partial Correlation Networks in Clinical Multi-omics Studies with HP-ACCORD
- 基于伪似然重构稀疏精度矩阵,用新损失函数+L1正则优化
- 百万变量模拟验证有效,真实肝癌数据中精准识别关键转录因子
- 适合高维多组学分析,尤其关注基因表达与表观调控分离的研究
从多组学数据构建图模型需兼顾统计性能与计算可扩展性。本文提出一种基于伪似然的新图模型框架,通过重参数化目标精度矩阵并保持稀疏结构,基于新损失函数和L1正则化最小化经验风险进行估计。该估计器在高维假设下具备估计与选择一致性。相关优化问题采用新型算子分裂方法和通信规避分布式矩阵乘法,实现可证明快速计算。框架在含百万变量的模拟数据上测试,成功捕获类似生物网络的复杂依赖结构。利用其可扩展性,我们基于双组学肝癌数据集构建部分相关网络。该共表达网络在超高维数据中表现出更优特异性,排除了表观遗传调控干扰,精准识别关键转录因子与共激活因子,凸显计算可扩展性在多组学分析中的价值。
原文摘要 · Abstract (English)
Graphical model estimation from multi-omics data requires a balance between statistical estimation performance and computational scalability. We introduce a novel pseudolikelihood-based graphical model framework that reparameterizes the target precision matrix while preserving the sparsity pattern and estimates it by minimizing an $\ell_1$-penalized empirical risk based on a new loss function. The proposed estimator maintains estimation and selection consistency in various metrics under high-dimensional assumptions. The associated optimization problem allows for a provably fast computation algorithm using a novel operator-splitting approach and communication-avoiding distributed matrix multiplication. A high-performance computing implementation of our framework was tested using simulated data with up to one million variables, demonstrating complex dependency structures similar to those found in biological networks. Leveraging this scalability, we estimated a partial correlation network from a dual-omic liver cancer data set. The co-expression network estimated from the ultrahigh-dimensional data demonstrated superior specificity in prioritizing key transcription factors and co-activators by excluding the impact of epigenetic regulation, thereby highlighting the value of computational scalability in multi-omic data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。