让蛋白语言模型像智能代理一样边生成边纠错,提升设计成功率。
AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

- 用工具调用与自回归生成交替进行,动态评估结构可行性。
- 抗体设计任务中顶尖表现,较最强基线提升10%命中率。
- 适合蛋白质设计、药物开发等需要高精度生成的场景。
蛋白质语言模型(PLMs)是被动的生成器:它们单次前向传播生成序列,缺乏机制来获取外部生物物理反馈或在候选序列违反热力学/结构约束时调整生成路径。我们提出AgentPLM,通过两个核心机制解决此问题:一、推理增强解码(RAD),将自回归生成与工具调用(ESMFold、FoldX、AutoDock Vina)交替进行;二、对比式代理策略优化(CAPO),作为直接偏好优化的轨迹级扩展,端到端训练策略以判断何时利用生成反馈而非简单模仿高适应度序列。我们在涵盖从头酶设计、抗体优化、热稳定性、蛋白质-蛋白质相互作用界面设计及零样本适应度预测的基准任务上评估该方法,使用标准化的外部评估接口和受控序列相似性划分。AgentPLM在各项任务中达到当前最优性能,在抗体优化任务中相比最强被动基线提升顶10%命中率,提供了无需显式回溯即可实现在线错误纠正的机制证据。
原文摘要 · Abstract (English)
Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural constraints. We introduce AgentPLM, which addresses this by equipping a pre-trained PLM with i) Reasoning-Augmented Decoding (RAD), which interleaves autoregressive generation with tool calls (ESMFold, FoldX, AutoDock Vina), and ii) Contrastive Agent Policy Optimisation (CAPO), a trajectory-level extension of direct preference optimisation that trains the policy end-to-end to learn when oracle feedback is informative rather than merely imitating high-fitness sequences. We evaluate AgentPLM on benchmark tasks spanning de novo enzyme design, antibody optimisation, thermostability, PPI interface design, and zero-shot fitness prediction with standardised oracle APIs and controlled sequence-identity splits. AgentPLM achieves state-of-the-art results with a gain in antibody top-10% hit rate over the strongest passive baseline, providing mechanistic evidence of online error correction without explicit backtracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。