构建6个超大规模单细胞基础模型,显著提升疾病机制理解能力
TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology
- 用1.16亿细胞数据+生物注释监督预训练,提升模型泛化性
- 在未见患者和疾病状态下,识别准确率大幅超越现有模型
- 适合生物医学研究者、药物研发人员快速获取细胞机制洞察
理解疾病生物学机制对医学和药物发现至关重要。人工智能分析基因组规模生物数据在此领域潜力巨大。单细胞RNA测序数据的日益丰富,推动了疾病生物学大基础模型的发展。然而,现有基础模型在下游任务中仅小幅优于专用模型。本文探索两条优化路径:一是将预训练数据量扩大至1.16亿细胞(超过以往模型),二是利用大规模生物注释作为预训练阶段的监督信号。我们训练了包含六个Transformer架构的模型家族,参数量分别为7000万、1.6亿和4亿。在多个下游评估任务中验证性能,包括对未见供体的疾病状态识别、未见患者与疾病条件下的健康/病变细胞区分,以及对学习表征中已知生物学知识的探查。结果表明,模型显著优于现有工作;缩放实验显示,性能随数据量和参数量增加而可预测地提升。
原文摘要 · Abstract (English)
Understanding the biological mechanisms of disease is crucial for medicine, and in particular, for drug discovery. AI-powered analysis of genome-scale biological data holds great potential in this regard. The increasing availability of single-cell RNA sequencing data has enabled the development of large foundation models for disease biology. However, existing foundation models only modestly improve over task-specific models in downstream applications. Here, we explored two avenues for improving single-cell foundation models. First, we scaled the pre-training data to a diverse collection of 116 million cells, which is larger than those used by previous models. Second, we leveraged the availability of large-scale biological annotations as a form of supervision during pre-training. We trained the \model family of models comprising six transformer-based state-of-the-art single-cell foundation models with 70 million, 160 million, and 400 million parameters. We vetted our models on several downstream evaluation tasks, including identifying the underlying disease state of held-out donors not seen during training, distinguishing between diseased and healthy cells for disease conditions and donors not seen during training, and probing the learned representations for known biology. Our models showed substantial improvement over existing works, and scaling experiments showed that performance improved predictably with both data volume and parameter count.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。