arXiv:2412.13478cs.LGq-bio.QM2024-12被引 14

用少量参数微调单细胞模型,实现新药反应的零样本预测

Efficient Fine-Tuning of Single-Cell Foundation Models Enables Zero-Shot Molecular Perturbation Prediction

  • 设计药物条件适配器,仅训练不足1%参数实现高效微调
  • 在新细胞系上零样本预测准确率显著优于现有方法
  • 适用于新药研发、罕见疾病研究等需要快速推断的场景

预测新药物引起的转录响应为加速生物医学研究和药物发现提供了独特机遇。然而,细胞响应的固有复杂性和高维性,以及实验数据极度稀缺,使该任务极具挑战。本研究利用在数百万个单细胞(涵盖多种细胞类型、状态及疾病注释)上预训练的单细胞基础模型,解决分子扰动预测问题。我们提出一种药物条件适配器,通过训练原模型不到1%的参数,实现高效微调,在保持预训练中学习到的丰富生物学表征的同时,实现分子条件化。该策略不仅能预测新药的细胞响应,还能在未见过的细胞系上实现零样本泛化。我们建立了稳健的评估框架,评估不同泛化任务下的模型表现,结果在所有设置下均达到当前最优,尤其在少样本和零样本跨细胞系泛化方面相比基线有显著提升。

原文摘要 · Abstract (English)

Predicting transcriptional responses to novel drugs provides a unique opportunity to accelerate biomedical research and advance drug discovery efforts. However, the inherent complexity and high dimensionality of cellular responses, combined with the extremely limited available experimental data, makes the task challenging. In this study, we leverage single-cell foundation models (FMs) pre-trained on tens of millions of single cells, encompassing multiple cell types, states, and disease annotations, to address molecular perturbation prediction. We introduce a drug-conditional adapter that allows efficient fine-tuning by training less than 1% of the original foundation model, thus enabling molecular conditioning while preserving the rich biological representation learned during pre-training. The proposed strategy allows not only the prediction of cellular responses to novel drugs, but also the zero-shot generalization to unseen cell lines. We establish a robust evaluation framework to assess model performance across different generalization tasks, demonstrating state-of-the-art results across all settings, with significant improvements in the few-shot and zero-shot generalization to new cell lines compared to existing baselines.

单细胞零样本药物发现微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。