融合两种蛋白质组学数据,提升胰腺癌分类精度
Learning predictive models for combinations of heterogeneous proteomic data sources
- 提出模型融合方法,应对不同数据源间的差异性
- 单一模型在组合数据上表现下降,融合后准确率显著提升
- 适合多组学数据整合与生物标志物发现的研究者
多种测量人体蛋白混合物表达水平的技术为疾病检测与理解提供了可能。近期技术的增加促使研究者评估各技术单独及联合使用的价值。本文研究了两种数据源:全样本质谱分析和多重蛋白阵列。通过在胰腺癌研究数据上训练和测试多种分类模型,发现针对单一数据源表现良好的模型在二者组合时性能下降。为此,本文提出一类模型融合方法,充分考虑数据源差异,最大化联合数据的利用效益。
原文摘要 · Abstract (English)
Multiple technologies that measure expression levels of protein mixtures in the human body offer a potential for detection and understanding the disease. The recent increase of these technologies prompts researchers to evaluate the individual and combined utility of data generated by the technologies. In this work, we study two data sources to measure the expression of protein mixtures in the human body: whole-sample MS profiling and multiplexed protein arrays. We investigate the individual and combined utility of these technologies by learning and testing a variety of classification models on the data from a pancreatic cancer study. We show that for the combination of these two (heterogeneous) datasets, classification models that work well on one of them individually fail on the combination of the two datasets. We study and propose a class of model fusion methods that acknowledge the differences and try to reap most of the benefits from their combination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。