用生成预训练+测试时计算,统一蛋白结合剂设计的生成与优化方法。
Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
- 融合生成模型与测试时优化,实现原子级精确的蛋白结合剂设计。
- 在模拟测试中成功率显著高于现有生成方法,计算效率提升超2倍。
- 适用于抗体、小分子靶点及酶设计,适合药物研发与蛋白质工程人群。
蛋白质相互作用建模是蛋白质设计的核心,机器学习已推动其在药物发现等领域的变革。当前结构引导的从头结合剂设计主要分为条件生成建模或通过结构预测器进行序列优化(即“幻觉”)。我们指出这并非非此即彼的选择,提出 Proteina-Complexa——一种统一两种范式的全原子级结合剂生成新方法。该方法扩展了基于流的潜空间蛋白生成架构,并利用单体蛋白结构的域间相互作用,构建了 Teddymer 这一大规模合成结合对数据集用于预训练。结合高质量实验多聚体数据,训练出强大基础模型。随后在推理阶段引入优化策略,融合生成先验与动态调整优势,显著优于以往独立方法。Proteina-Complexa 在多个计算结合剂设计基准上达到新最优:相比现有生成方法,仿真成功率大幅提升;测试时优化策略在标准化算力下性能超越先前幻觉方法超2倍。此外,还实现了界面氢键优化、折叠类别引导生成,并拓展至小分子靶点与酶设计任务,均优于已有方法。代码、模型与新数据将公开发布。
原文摘要 · Abstract (English)
Protein interaction modeling is central to protein design, which has been transformed by machine learning with applications in drug discovery and beyond. In this landscape, structure-based de novo binder design is cast as either conditional generative modeling or sequence optimization via structure predictors ("hallucination"). We argue that this is a false dichotomy and propose Proteina-Complexa, a novel fully atomistic binder generation method unifying both paradigms. We extend recent flow-based latent protein generation architectures and leverage the domain-domain interactions of monomeric computationally predicted protein structures to construct Teddymer, a new large-scale dataset of synthetic binder-target pairs for pretraining. Combined with high-quality experimental multimers, this enables training a strong base model. We then perform inference-time optimization with this generative prior, unifying the strengths of previously distinct generative and hallucination methods. Proteina-Complexa sets a new state of the art in computational binder design benchmarks: it delivers markedly higher in-silico success rates than existing generative approaches, and our novel test-time optimization strategies greatly outperform previous hallucination methods under normalized compute budgets. We also demonstrate interface hydrogen bond optimization, fold class-guided binder generation, and extensions to small molecule targets and enzyme design tasks, again surpassing prior methods. Code, models and new data will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。