用集成学习融合多组学数据,提升肝癌早期诊断准确率。
Multi-omics data integration for early diagnosis of hepatocellular carcinoma (HCC) using machine learning
- 采用多种集成模型实现多模态数据晚期融合
- 最高准确率达AUC 0.85,PB-MVBoost与软投票Adaboost最优
- 适合生物标志物发现与临床决策支持研究者参考
不同患者数据模态间具有互补信息,有助于更准确地建模疾病状态并理解其生物学机制。然而,多模态、多组学数据的分析面临高维性、模态间规模、分布、尺度和信号强度差异等挑战。本文比较了多种可实现多类数据晚期融合的集成学习算法:包括硬/软投票集成、元学习器、基于硬/软投票及元学习器的多模态Adaboost(PB-MVBoost),以及混合专家模型。以肝细胞癌(HCC)的内部数据及乳腺癌、肠易激综合征(IBD)四个验证数据集进行评估。以受试者工作特征曲线下面积(AUC)为评价指标,所建模型最高达到AUC 0.85;其中,PB-MVBoost与软投票Adaboost表现最佳。还分析了特征选择稳定性与临床标志物规模,并对多模态多分类数据整合提出建议。
原文摘要 · Abstract (English)
The complementary information found in different modalities of patient data can aid in more accurate modelling of a patient's disease state and a better understanding of the underlying biological processes of a disease. However, the analysis of multi-modal, multi-omics data presents many challenges, including high dimensionality and varying size, statistical distribution, scale and signal strength between modalities. In this work we compare the performance of a variety of ensemble machine learning algorithms that are capable of late integration of multi-class data from different modalities. The ensemble methods and their variations tested were i) a voting ensemble, with hard and soft vote, ii) a meta learner, iii) a multi-modal Adaboost model using a hard vote, a soft vote and a meta learner to integrate the modalities on each boosting round, the PB-MVBoost model and a novel application of a mixture of experts model. These were compared to simple concatenation as a baseline. We examine these methods using data from an in-house study on hepatocellular carcinoma (HCC), along with four validation datasets on studies from breast cancer and irritable bowel disease (IBD). Using the area under the receiver operating curve as a measure of performance we develop models that achieve a performance value of up to 0.85 and find that two boosted methods, PB-MVBoost and Adaboost with a soft vote were the overall best performing models. We also examine the stability of features selected, and the size of the clinical signature determined. Finally, we provide recommendations for the integration of multi-modal multi-class data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。