arXiv:2411.15180cs.LGcs.AI2024-11被引 2

用多层分解融合多组学数据,提升癌症亚型分类准确率

Multi-layer matrix factorization for cancer subtyping using full and partial multi-omics dataset

  • 通过多层矩阵分解提取各组学类型的潜在特征
  • 在10个数据集上表现优于或媲美现有方法
  • 可处理完整与缺失数据,适合真实临床场景

癌症因其内在异质性常根据独特特征、细胞起源和分子标志物分为不同亚型。然而,现有研究多依赖完整的多组学数据进行亚型预测,忽视了部分数据缺失时的性能表现,也未充分挖掘多层组学数据间的隐含关联。本文提出多层矩阵分解(MLMF)方法,通过多层线性或非线性分解将多组学特征矩阵转化为每类组学独有的潜在表示,并融合为共识形式后进行谱聚类以确定亚型。此外,MLMF引入类别指示矩阵处理缺失组学数据,构建统一框架,可同时处理完整与不完整数据。在10个包含完整与缺失值的多组学癌症数据集上的大量实验表明,MLMF性能达到或超过多个先进方法。

原文摘要 · Abstract (English)

Cancer, with its inherent heterogeneity, is commonly categorized into distinct subtypes based on unique traits, cellular origins, and molecular markers specific to each type. However, current studies primarily rely on complete multi-omics datasets for predicting cancer subtypes, often overlooking predictive performance in cases where some omics data may be missing and neglecting implicit relationships across multiple layers of omics data integration. This paper introduces Multi-Layer Matrix Factorization (MLMF), a novel approach for cancer subtyping that employs multi-omics data clustering. MLMF initially processes multi-omics feature matrices by performing multi-layer linear or nonlinear factorization, decomposing the original data into latent feature representations unique to each omics type. These latent representations are subsequently fused into a consensus form, on which spectral clustering is performed to determine subtypes. Additionally, MLMF incorporates a class indicator matrix to handle missing omics data, creating a unified framework that can manage both complete and incomplete multi-omics data. Extensive experiments conducted on 10 multi-omics cancer datasets, both complete and with missing values, demonstrate that MLMF achieves results that are comparable to or surpass the performance of several state-of-the-art approaches.

癌症亚型多组学分析矩阵分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。