arXiv:2503.23535cs.LGq-bio.QM2025-03被引 19

用深度学习统一整合多种扰动实验数据,加速生物发现

In-silico biological discovery with large perturbation models

  • 将扰动、读出和生物背景解耦建模,统一处理异构实验数据
  • 在未见实验中预测转录组变化,准确率显著优于现有方法
  • 适合药物机制解析与基因互作网络推断的研究者使用

扰动实验数据揭示了扰动与其生物学效应之间的关联,对理解生物实体关系和开发治疗手段具有重要意义。然而,这些数据涉及多种扰动方式与检测指标,且实验结果受生物背景复杂影响,难以跨实验整合。本文提出大型扰动模型(LPM),通过将扰动、读出和上下文分别表示为解耦维度,实现对多源异构扰动实验数据的统一建模。LPM在多个生物发现任务中表现优异,包括预测未知实验的扰后转录组、识别化学与遗传扰动的共享作用机制,以及推断基因-基因互作网络。

原文摘要 · Abstract (English)

Data generated in perturbation experiments link perturbations to the changes they elicit and therefore contain information relevant to numerous biological discovery tasks -- from understanding the relationships between biological entities to developing therapeutics. However, these data encompass diverse perturbations and readouts, and the complex dependence of experimental outcomes on their biological context makes it challenging to integrate insights across experiments. Here, we present the Large Perturbation Model (LPM), a deep-learning model that integrates multiple, heterogeneous perturbation experiments by representing perturbation, readout, and context as disentangled dimensions. LPM outperforms existing methods across multiple biological discovery tasks, including in predicting post-perturbation transcriptomes of unseen experiments, identifying shared molecular mechanisms of action between chemical and genetic perturbations, and facilitating the inference of gene-gene interaction networks.

生物发现深度学习扰动模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。