用半监督学习提升癌症预后预测准确率
Deep Semi-Supervised Survival Analysis for Predicting Cancer Prognosis
- 基于Mean Teacher框架,融合标签与无标签数据训练深度Cox模型
- 在四种癌症上表现优于现有模型,无标签数据越多效果越显著
- 适合处理标注数据稀缺的医疗生存分析任务
Cox比例风险模型广泛应用于生存分析。近年来,基于人工神经网络(ANN)的Cox-PH模型被提出。然而,使用高维特征训练这些Cox模型通常需要大量包含事件时间信息的标注样本,而标注数据有限常制约模型性能。为解决此问题,本文采用深度半监督学习(DSSL)方法,在Mean Teacher(MT)框架下构建单模态与多模态的ANN-Cox模型,利用标签与无标签数据联合训练。所提模型命名为Cox-MT,应用于泛癌种数据集The Cancer Genome Atlas(TCGA),基于RNA-seq数据或全切片图像的单模态Cox-MT模型在四种癌症上均显著优于同数据集下的现有模型Cox-nnet。随着无标签样本数量增加,在固定标签数据量下,Cox-MT性能显著提升。此外,多模态Cox-MT模型表现明显优于单模态模型。结果表明,相较于仅依赖标签数据训练的现有模型,Cox-MT有效利用未标注数据,显著提升了预测精度。
原文摘要 · Abstract (English)
The Cox Proportional Hazards (PH) model is widely used in survival analysis. Recently, artificial neural network (ANN)-based Cox-PH models have been developed. However, training these Cox models with high-dimensional features typically requires a substantial number of labeled samples containing information about time-to-event. The limited availability of labeled data for training often constrains the performance of ANN-based Cox models. To address this issue, we employed a deep semi-supervised learning (DSSL) approach to develop single- and multi-modal ANN-based Cox models based on the Mean Teacher (MT) framework, which utilizes both labeled and unlabeled data for training. We applied our model, named Cox-MT, to predict the prognosis of several types of cancer using data from The Cancer Genome Atlas (TCGA). Our single-modal Cox-MT models, utilizing TCGA RNA-seq data or whole slide images, significantly outperformed the existing ANN-based Cox model, Cox-nnet, using the same data set across four types of cancer considered. As the number of unlabeled samples increased, the performance of Cox-MT significantly improved with a given set of labeled data. Furthermore, our multi-modal Cox-MT model demonstrated considerably better performance than the single-modal model. In summary, the Cox-MT model effectively leverages both labeled and unlabeled data to significantly enhance prediction accuracy compared to existing ANN-based Cox models trained solely on labeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。