构建首个单细胞自监督学习综合评测平台,揭示不同方法在数据整合中的优劣。
scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell Data
- 评测19种自监督学习方法在九个数据集上的表现,覆盖三种下游任务。
- 随机掩码增广优于领域特定策略,通用模型在细胞注释中表现更佳。
- 适合生物信息学研究者与深度学习开发者参考方法选型。
自监督学习(SSL)在从单细胞数据中提取生物学有意义表征方面已证明具有强大能力。为推进对单细胞数据中应用的SSL方法的理解,我们提出了scSSL-Bench,一个全面的基准测试平台,评估了19种SSL方法。评估涵盖九个数据集,并聚焦于三种常见下游任务:批次校正、细胞类型注释和缺失模态预测。此外,我们系统地评估了多种数据增强策略。分析揭示了任务相关的权衡:专用单细胞框架scVI、CLAIRE以及微调后的scGPT在单模态批次校正中表现优异,而通用SSL方法如VICReg和SimCLR在细胞类型分类和多模态数据整合中表现更优。随机掩码成为所有任务中最有效的增强技术,优于领域特定的增强方法。值得注意的是,我们的结果表明需要专门的单细胞多模态数据整合框架。scSSL-Bench提供了一个标准化的评估平台,并为将SSL应用于单细胞分析提供了具体建议,推动深度学习与单细胞基因组学的融合。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has proven to be a powerful approach for extracting biologically meaningful representations from single-cell data. To advance our understanding of SSL methods applied to single-cell data, we present scSSL-Bench, a comprehensive benchmark that evaluates nineteen SSL methods. Our evaluation spans nine datasets and focuses on three common downstream tasks: batch correction, cell type annotation, and missing modality prediction. Furthermore, we systematically assess various data augmentation strategies. Our analysis reveals task-specific trade-offs: the specialized single-cell frameworks, scVI, CLAIRE, and the finetuned scGPT excel at uni-modal batch correction, while generic SSL methods, such as VICReg and SimCLR, demonstrate superior performance in cell typing and multi-modal data integration. Random masking emerges as the most effective augmentation technique across all tasks, surpassing domain-specific augmentations. Notably, our results indicate the need for a specialized single-cell multi-modal data integration framework. scSSL-Bench provides a standardized evaluation platform and concrete recommendations for applying SSL to single-cell analysis, advancing the convergence of deep learning and single-cell genomics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。