MapPFN通过上下文学习实现单细胞扰动的自适应预测,无需重新训练。
MapPFN: Learning Causal Perturbation Maps in Context
- 用合成数据预训练的因果扰动网络,实现零样本泛化
- 在未见基因集上表现接近真实数据训练模型,微调后更优
- 适合需快速适配新实验数据的研究者,尤其生物干预场景
在生物系统中规划有效干预需要能适应未知生物背景的治疗效应模型,以识别其特定作用机制。然而,单细胞扰动数据集仅覆盖少量生物背景,现有方法无法在推理时利用新干预证据进行扩展。为此,我们提出MapPFN,一种基于合成生物学先验的前训练网络(PFN),将预训练与有限湿实验数据解耦。不同于以往方法,MapPFN采用上下文学习,将一系列实验映射为扰动后分布,使单一预训练模型可在推理时适配新数据集和任意基因集。零样本情况下,MapPFN识别差异表达基因的效果与真实数据训练模型相当;微调后,在多个生物背景下预测性能进一步提升。代码、模型与数据已公开于https://marvinsxtr.github.io/MapPFN。
原文摘要 · Abstract (English)
Planning effective interventions in biological systems requires treatment-effect models that adapt to unseen biological contexts by identifying their specific underlying mechanisms. Yet single-cell perturbation datasets span only a handful of biological contexts, and existing methods cannot leverage new interventional evidence at inference time to adapt beyond their training data. To meta-learn a perturbation effect estimator, we present MapPFN, a prior-data fitted network (PFN) pre-trained on a synthetic biological prior with causal interventions, decoupling pre-training from limited wet-lab data. Unlike existing methods, MapPFN uses in-context learning to map a sequence of experiments to a post-perturbation distribution, enabling a single pre-trained model to adapt to new datasets and arbitrary gene sets at inference time. Zero-shot, MapPFN identifies differentially expressed genes on par with models trained on real single-cell data, and fine-tuning further improves predictions across biological contexts. Our code, model and data are available at https://marvinsxtr.github.io/MapPFN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。