首个针对单细胞药物反应预测的大规模模型基准测试平台
scDrugMap: Benchmarking Large Foundation Models for Drug Response Prediction
- 构建scDrugMap框架,集成命令行与网页端,支持多种大模型评估
- 在超32万细胞数据上验证,最佳模型F1达0.971(冻结层)
- 适合药物发现与转化研究者使用,提供零样本与微调双模式
药物耐药性是癌症治疗的重大挑战。单细胞分析可揭示细胞异质性,但大规模基础模型在单细胞药物反应预测中的应用仍不充分。为此,我们开发了scDrugMap,一个集成命令行与网页服务的药物反应预测框架。该框架评估了包括8个单细胞模型和2个大语言模型在内的多种基础模型,基于包含超过32.6万细胞的主数据集和1.88万细胞的验证集,覆盖36个数据集及多种组织与癌症类型。我们在合并数据与跨数据集两种评估设置下进行了性能对比,采用层冻结与低秩适应(LoRA)微调策略。在合并数据场景中,scFoundation表现最佳,冻结层时平均F1为0.971,微调后为0.947,优于最差模型超50%。在跨数据集设置中,微调后UCE表现最优(平均F1: 0.774),scGPT在零样本学习中领先(平均F1: 0.858)。scDrugMap首次实现单细胞药物反应预测的大规模模型基准测试,为药物研发与转化研究提供用户友好、灵活的平台。
原文摘要 · Abstract (English)
Drug resistance presents a major challenge in cancer therapy. Single cell profiling offers insights into cellular heterogeneity, yet the application of large-scale foundation models for predicting drug response in single cell data remains underexplored. To address this, we developed scDrugMap, an integrated framework featuring both a Python command-line interface and a web server for drug response prediction. scDrugMap evaluates a wide range of foundation models, including eight single-cell models and two large language models, using a curated dataset of over 326,000 cells in the primary collection and 18,800 cells in the validation set, spanning 36 datasets and diverse tissue and cancer types. We benchmarked model performance under pooled-data and cross-data evaluation settings, employing both layer freezing and Low-Rank Adaptation (LoRA) fine-tuning strategies. In the pooled-data scenario, scFoundation achieved the best performance, with mean F1 scores of 0.971 (layer freezing) and 0.947 (fine-tuning), outperforming the lowest-performing model by over 50%. In the cross-data setting, UCE excelled post fine-tuning (mean F1: 0.774), while scGPT led in zero-shot learning (mean F1: 0.858). Overall, scDrugMap provides the first large-scale benchmark of foundation models for drug response prediction in single-cell data and serves as a user-friendly, flexible platform for advancing drug discovery and translational research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。