arXiv:2506.07459cs.LGq-bio.QM2025-06被引 10

用在线强化学习让蛋白生成模型自我进化,成功率超90%。

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

  • 通过自研能量预测器与ESMFold结合,实现低成本多目标反馈。
  • 在CATH-4.3上失败率降低36%-48%,成功率超90%。
  • 单机8卡3天完成一轮训练,适合持续优化蛋白设计。

蛋白质生成模型在设计中展现出巨大潜力,但受限于依赖人工标注的序列-结构数据集,以及监督目标与实际设计目标之间的错位。我们提出ProteinZero,一种针对逆折叠模型的在线强化学习框架,支持可扩展、自动化且持续的自我改进,采用计算高效的反馈机制。该框架融合ESMFold的结构引导与新型自衍生ddG预测器,提供稳定多目标信号,避免高成本物理方法。为保障在线强化学习的鲁棒性,引入嵌入层多样性正则化,缓解模式崩溃并促进功能有意义的序列变异。在兼顾多奖励优化、参考模型KL散度和多样性正则化的通用强化学习范式下,ProteinZero在可设计性、稳定性、恢复率和多样性上均实现稳健提升。在CATH-4.3基准测试中,其性能持续优于ProteinMPNN、ESM-IF和InstructPLM等先进基线,设计失败率降低36%-48%,各类折叠的成功率均超过90%。值得注意的是,完整强化学习流程可在单台8卡节点上3日内完成,包括奖励计算与数据生成。结果表明,高效在线强化学习微调可补充监督预训练,使蛋白生成模型基于自身输出持续演化,无需标签数据即可优化多个设计目标,为探索广阔蛋白设计空间开辟新路径。发布后将开放全部源代码与模型检查点。

原文摘要 · Abstract (English)

Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals. We present ProteinZero, an online reinforcement learning framework for inverse folding models that enables scalable, automated, and continuous self-improvement with computationally efficient feedback. ProteinZero employs a reward pipeline that combines structural guidance from ESMFold with a novel self-derived ddG predictor, providing stable multi-objective signals while avoiding the prohibitive cost of physics-based methods. To ensure robustness in online RL, we further introduce a novel embedding-level diversity regularizer that mitigates mode collapse and promotes functionally meaningful sequence variation. Within a general RL formulation balancing multi-reward optimization, KL-divergence from a reference model, and diversity regularization, ProteinZero achieves robust improvements across designability, stability, recovery, and diversity. On the CATH-4.3 benchmark, it consistently outperforms state-of-the-art baselines including ProteinMPNN, ESM-IF, and InstructPLM, reducing design failure rates by 36-48% and achieving success rates above 90% across diverse folds. Importantly, a complete RL run can be executed on a single 8 X GPU node within three days, including reward computation and data generation. These results indicate that efficient online RL fine-tuning can complement supervised pretraining by allowing protein generative models to evolve continuously from their own outputs and optimize multiple design objectives without labeled data, opening new possibilities for exploring the vast protein design space. Full source code and model checkpoints will be released upon publication.

蛋白生成强化学习自进化结构设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。