用自监督学习从基因表达数据中提取特征,减少对标注数据依赖
Self-supervised learning on gene expression data
- 采用三种不同原理的自监督方法,从无标签表达数据中学习特征
- 在多个公开数据集上提升表型预测准确率,优于传统监督模型
- 适合缺乏标注数据的生物医学研究,为个性化医疗提供新思路
从基因表达数据预测表型是生物医学研究的关键任务,有助于揭示疾病机制、药物反应和个性化医疗。传统机器学习与深度学习依赖大量标注数据,而基因表达数据的标注成本高、耗时长。自监督学习通过直接挖掘无标签数据的内在结构,成为克服这一瓶颈的有前景方法。本研究首次系统评估了前沿自监督学习方法在批量基因表达数据(bulk RNA-Seq)中的应用,选取三种基于不同机制的方法,分析其捕捉复杂信息并生成可用于下游预测的高质量表示的能力。基于多个公开基因表达数据集的实验表明,所选方法能有效利用数据结构,显著提升表型预测性能,且相比传统监督模型具有更少依赖标注数据的优势。研究还深入对比各方法优劣,提出适用建议,并展望未来发展方向。该工作为自监督学习在基因表达分析中的应用提供了首个全面实证基础。
原文摘要 · Abstract (English)
Predicting phenotypes from gene expression data is a crucial task in biomedical research, enabling insights into disease mechanisms, drug responses, and personalized medicine. Traditional machine learning and deep learning rely on supervised learning, which requires large quantities of labeled data that are costly and time-consuming to obtain in the case of gene expression data. Self-supervised learning has recently emerged as a promising approach to overcome these limitations by extracting information directly from the structure of unlabeled data. In this study, we investigate the application of state-of-the-art self-supervised learning methods to bulk gene expression data for phenotype prediction. We selected three self-supervised methods, based on different approaches, to assess their ability to exploit the inherent structure of the data and to generate qualitative representations which can be used for downstream predictive tasks. By using several publicly available gene expression datasets, we demonstrate how the selected methods can effectively capture complex information and improve phenotype prediction accuracy. The results obtained show that self-supervised learning methods can outperform traditional supervised models besides offering significant advantage by reducing the dependency on annotated data. We provide a comprehensive analysis of the performance of each method by highlighting their strengths and limitations. We also provide recommendations for using these methods depending on the case under study. Finally, we outline future research directions to enhance the application of self-supervised learning in the field of gene expression data analysis. This study is the first work that deals with bulk RNA-Seq data and self-supervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。