arXiv:2510.24670cs.LGq-bio.QM2025-10被引 8

珍珠模型精准预测蛋白-配体复合物三维结构,提升药物设计效率。

Pearl: A Foundation Model for Placing Every Atom in the Right Location

  • 基于大规模合成数据与旋转对称性保持的扩散模块,提升泛化能力。
  • 在多个基准上达到新纪录,比次优模型提升14.5%以上(RMSD<2Å)。
  • 支持条件与非条件推理,适合真实药物靶点复杂场景。

准确预测蛋白-配体复合物的三维结构仍是计算药物发现中的核心挑战,制约治疗设计的速度与成功率。深度学习方法虽展现强大潜力,但受限于实验数据稀缺、架构低效、物理无效构象及推理时难以利用辅助信息。为此,我们提出珍珠(Pearl:Placing Every Atom in the Right Location)——一种可规模化执行蛋白-配体共折叠的基础模型。珍珠通过三项关键创新解决上述问题:(1) 采用包含大规模合成数据的训练方案,缓解数据稀缺;(2) 引入SO(3)-等变扩散模块,天然保留三维旋转对称性,提升泛化与采样效率;(3) 可控推理机制,包括支持蛋白质与非聚合组分的多链模板系统,以及无条件/有条件双模式。珍珠在蛋白-配体共折叠任务中建立新基准。在生成准确且物理有效的构象(RMSD < 2 Å)的关键指标上,其在公开的Runs N' Poses和PoseBusters基准上分别优于AlphaFold 3及其他开源基线模型14.5%和14.2%。在口袋条件共折叠场景下,针对具有挑战性的私有真实药物靶点,在更严格的RMSD < 1 Å阈值下实现3.6倍性能提升。最后,我们验证了模型性能与训练所用合成数据规模直接相关。

原文摘要 · Abstract (English)

Accurately predicting the three-dimensional structures of protein-ligand complexes remains a fundamental challenge in computational drug discovery that limits the pace and success of therapeutic design. Deep learning methods have recently shown strong potential as structural prediction tools, achieving promising accuracy across diverse biomolecular systems. However, their performance and utility are constrained by scarce experimental data, inefficient architectures, physically invalid poses, and the limited ability to exploit auxiliary information available at inference. To address these issues, we introduce Pearl (Placing Every Atom in the Right Location), a foundation model for protein-ligand cofolding at scale. Pearl addresses these challenges with three key innovations: (1) training recipes that include large-scale synthetic data to overcome data scarcity; (2) architectures that incorporate an SO(3)-equivariant diffusion module to inherently respect 3D rotational symmetries, improving generalization and sample efficiency, and (3) controllable inference, including a generalized multi-chain templating system supporting both protein and non-polymeric components as well as dual unconditional/conditional modes. Pearl establishes a new state-of-the-art performance in protein-ligand cofolding. On the key metric of generating accurate (RMSD < 2 Å) and physically valid poses, Pearl surpasses AlphaFold 3 and other open source baselines on the public Runs N' Poses and PoseBusters benchmarks, delivering 14.5% and 14.2% improvements, respectively, over the next best model. In the pocket-conditional cofolding regime, Pearl delivers $3.6\times$ improvement on a proprietary set of challenging, real-world drug targets at the more rigorous RMSD < 1 Å threshold. Finally, we demonstrate that model performance correlates directly with synthetic dataset size used in training.

蛋白结构药物设计扩散模型基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。