arXiv:2604.13986cs.LG2026-04被引 3

用流匹配建模基因扰动对细胞状态的影响,提升药物发现效率。

PRiMeFlow: Capturing Complex Expression Heterogeneity in Perturbation Response Modelling

  • 基于流匹配直接在基因表达空间建模扰动效应。
  • 在多个数据集上准确逼近单细胞表达分布,表现优于基准方法。
  • 适用于大规模扰动数据,适合生物医学研究与药物研发人员。

预测体外基因或小分子扰动对细胞状态的影响,可大规模识别细胞行为驱动因子并加速药物发现。然而,由于单细胞基因表达的固有异质性以及复杂的潜在基因依赖关系,建模仍面临挑战。本文提出PRiMeFlow,一种基于流匹配的端到端方法,直接在基因表达空间中建模遗传和小分子扰动的影响。通过在PerturBench中的广泛基准测试,我们证明了该方法能准确逼近单细胞基因表达的实证分布。消融实验验证了在基因表达空间操作、使用U-Net参数化速度场等关键设计选择的有效性。最后,通过将PRiMeFlow扩展至涵盖多个数据集的扰动数据图谱,并采用精心设计的预训练-微调策略,其在ARC虚拟细胞挑战赛中的人类胚胎干细胞(H1)数据集上展现出卓越性能。

原文摘要 · Abstract (English)

Predicting the effects of perturbations in-silico on cell state can identify drivers of cell behavior at scale and accelerate drug discovery. However, modeling challenges remain due to the inherent heterogeneity of single cell gene expression and the complex, latent gene dependencies. Here, we present PRiMeFlow, an end-to-end flow matching based approach to directly model the effects of genetic and small molecule perturbations in the gene expression space. The distribution-fitting approach taken by PRiMeFlow enables it to accurately approximate the empirical distribution of single-cell gene expression, which we demonstrate through extensive benchmarking inside PerturBench. Through ablation studies, we also validate important model design choices such as operating in gene expression space and parameterizing the velocity field with a U-Net architecture. Finally, by scaling PRiMeFlow to a broad perturbation data atlas spanning multiple datasets and employing a carefully designed pretraining-finetuning strategy, we demonstrate its outstanding performance on the H1 human embryonic stem cells from the ARC Virtual Cell Challenge benchmark.

基因扰动流匹配单细胞分析药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。