arXiv:2410.02496cs.LG2024-10

高效学习多源非正态图模型间的稀疏差异网络,提升推断速度与准确率。

Efficient learning of differential network in multi-source non-paranormal graphical models

  • 通过Lasso惩罚的D-Trace损失函数优化差分精度矩阵,解路径精确可算。
  • 在合成数据上速度与准确率均优于现有方法,尤其在稀疏差异网络下表现突出。
  • 适用于多源异构数据,实证揭示肿瘤耐药关键基因,具生物学意义。

本文研究两类非正态图模型之间的稀疏结构变化(即差分网络)学习问题。假设每类数据来自多个异质来源,且所有非正态图模型具有相同的协方差矩阵。差分网络由精度矩阵之差编码,可通过优化Lasso惩罚的D-Trace损失函数进行解码。为此提出一种高效算法,能输出精确解路径,优于以往仅在预设正则化参数下采样的方法。该方法计算复杂度低,尤其在差分网络稀疏时优势明显。合成数据实验表明,本方法在速度和准确率上均优于现有方法。真实场景实验显示,融合多源数据策略在肿瘤耐药性研究中有效,识别出已有多项独立研究证实的关键耐药基因。

原文摘要 · Abstract (English)

This paper addresses learning of sparse structural changes or differential network between two classes of non-paranormal graphical models. We assume a multi-source and heterogeneous dataset is available for each class, where the covariance matrices are identical for all non-paranormal graphical models. The differential network, which are encoded by the difference precision matrix, can then be decoded by optimizing a lasso penalized D-trace loss function. To this aim, an efficient approach is proposed that outputs the exact solution path, outperforming the previous methods that only sample from the solution path in pre-selected regularization parameters. Notably, our proposed method has low computational complexity, especially when the differential network are sparse. Our simulations on synthetic data demonstrate a superior performance for our strategy in terms of speed and accuracy compared to an existing method. Moreover, our strategy in combining datasets from multiple sources is shown to be very effective in inferring differential network in real-world problems. This is backed by our experimental results on drug resistance in tumor cancers. In the latter case, our strategy outputs important genes for drug resistance which are already confirmed by various independent studies.

图模型差分网络多源学习稀疏推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。