用合成先验预测未知药物效应,实现低开销可解释建模
PerturbPFN: Probing the Limits of Synthetic Priors in Drug Perturbation Modelling
- 基于层级合成结构先验,通过隐式图结构推断药物靶点与作用强度
- 在真实与合成数据上均表现良好,预测精度接近专用模型且推理成本低
- 适合需可解释性与快速响应的药物研发场景
由于未知靶点和作用机制、高维表达响应以及小分子设计空间实验覆盖有限,预测细胞对未见化学扰动的响应极具挑战。我们提出PerturbPFN,一种基于PFN架构的无目标扰动预测模型,采用分层合成结构先验。该模型不直接回归高维表达响应,而是推断隐含系统图、稀疏原子干预靶点及干预强度,并通过结构因果模型解码器传播影响。模型完全在生物启发的图与表达模拟器生成的合成训练样本上训练,实现无需测试时梯度更新的结构化上下文学习。我们在真实单细胞扰动数据和合成基准上评估了该模型,涵盖效应预测、靶点识别与调控结构发现。结果表明,PerturbPFN在保持低推理开销的同时,实现了与专用基线相当的扰动预测性能,并揭示了可解释的中间估计——包括靶点、强度与系统结构。
原文摘要 · Abstract (English)
Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space. We propose PerturbPFN, a PFN-style amortized model for unknown-target perturbation prediction under a hierarchical synthetic structural prior. Instead of directly regressing high-dimensional expression responses, PerturbPFN infers a latent system graph, sparse atomic intervention targets, and intervention strengths, then propagates their effects through an SCM decoder. The model is trained entirely on prior-predictive synthetic episodes generated from biologically motivated graph and expression simulators, enabling structured in-context learning without test-time gradient updates. We evaluate PerturbPFN on both real single-cell perturbation data and synthetic benchmarks, covering effect prediction, target identification, and regulatory structure discovery. Our results show that PerturbPFN offers a complementary trade-off to specialized baselines, achieving competitive perturbation prediction with low inference cost while exposing interpretable intermediate estimates of targets, strengths, and system structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。