arXiv:2510.01571cs.LGcs.AI2025-10被引 5

用强化学习探索蛋白质语言模型的隐藏能力,发现新设计规律。

From Supervision to Exploration: What Does Protein Language Model Learn During Reinforcement Learning?

  • 将强化学习与蛋白质语言模型结合,提升序列设计效率。
  • 在四个任务中均显著提高成功率和采样效率,最高提升3.2倍。
  • 揭示奖励精度、策略容量与任务空间是决定性能的关键因素。

蛋白质语言模型(PLMs)通过大规模预训练和可扩展架构推动了计算蛋白质科学的发展。与此同时,强化学习(RL)拓展了探索范围,并实现了蛋白质设计中的多目标精确优化。然而,强化学习能否突破预训练先验,揭示潜在的序列-结构-功能规律仍不明确。本文通过在抗菌肽设计、激酶变异体优化、抗体工程和逆折叠四个领域结合不同强化学习算法与模型类型,探究强化学习是否能提升采样效率并发现监督学习无法捕捉的能力。实验表明,强化学习在多个基准测试中持续提升成功率与样本效率。性能受三重交互影响:任务剩余空间、奖励保真度与策略容量共同决定增益。当奖励准确且信息丰富、策略容量充足、任务存在超越监督基线的空间时,性能提升显著;反之,即使探索充分,增益也会趋于饱和。该研究为蛋白质设计中的强化学习应用提供实践指导:优先优化奖励建模与校准,根据任务难度匹配算法与正则化强度,并在边际收益最大的位置分配计算资源。代码已开源:https://github.com/chq1155/RL-PLM。

原文摘要 · Abstract (English)

Protein language models (PLMs) have advanced computational protein science through large-scale pretraining and scalable architectures. In parallel, reinforcement learning (RL) has broadened exploration and enabled precise multi-objective optimization in protein design. Yet whether RL can push PLMs beyond their pretraining priors to uncover latent sequence-structure-function rules remains unclear. We address this by pairing RL with PLMs across four domains: antimicrobial peptide design, kinase variant optimization, antibody engineering, and inverse folding. Using diverse RL algorithms and model classes, we ask if RL improves sampling efficiency and, more importantly, if it reveals capabilities not captured by supervised learning. Across benchmarks, RL consistently boosts success rates and sample efficiency. Performance follows a three-factor interaction: task headroom, reward fidelity, and policy capacity jointly determine gains. When rewards are accurate and informative, policies have sufficient capacity, and tasks leave room beyond supervised baselines, improvements scale; when rewards are noisy or capacity is constrained, gains saturate despite exploration. This view yields practical guidance for RL in protein design: prioritize reward modeling and calibration before scaling policy size, match algorithm and regularization strength to task difficulty, and allocate capacity where marginal gains are largest. Implementation is available at https://github.com/chq1155/RL-PLM.

蛋白质设计强化学习语言模型生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。