用大模型在推理时优化蛋白质序列设计,提升结构精度与成功率。
RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

- 用大模型作为生成优化器,在计算预算内搜索最优序列。
- 修复主流模型漏掉的高保真序列,成功率提升2.5倍,精度提高18%-68%。
- 适用于多种模型和骨架,支持图像反馈的多模态设计,无需重训练。
我们提出 RosettaSearch,一种基于骨架条件的蛋白质序列设计推理时多目标优化方法。利用大语言模型(LLMs)作为生成优化器,在搜索算法中实现受控探索与利用,奖励由结构预测模型 RosettaFold3 计算,且严格控制在计算预算内。大规模评估中,将 RosettaSearch 应用于 400 条 LigandMPNN 生成的次优序列,恢复出 LigandMPNN 单次解码无法产出的高保真设计。其设计在结构保真度指标上提升 18% 至 68%,设计成功率提升 2.5 倍。该成功率增益在独立结构预测验证器 Chai-1 下依然稳健,并在 o4-mini 与 Gemini-3 两类大模型上均表现一致,性能随推理能力提升而增长。此外,该方法还显著提升 ProteinMPNN 在 Dayhoff atlas 生成的从头骨架上的序列保真度,表明其可推广至计算生成骨架。我们进一步展示了结合视觉-语言模型的多模态扩展,以预测蛋白结构图像为反馈,引入结构上下文指导序列生成。据我们所知,这是首次大规模证明大模型可在不重训练的前提下,有效充当骨架条件蛋白序列设计的生成优化器并带来系统性提升。
原文摘要 · Abstract (English)
We introduce RosettaSearch, an inference-time multi-objective optimization approach for backbone conditioned protein sequence design. We use large language models (LLMs) as a generative optimizer within a search algorithm capable of controlled exploration and exploitation, using rewards computed from RosettaFold3, a structure prediction model, under a strict computational budget. In a large-scale evaluation, we apply RosettaSearch to 400 suboptimal sequences generated by LigandMPNN (a state-of-the-art model trained for protein sequence design), recovering high-fidelity designs that LigandMPNN's single-pass decoding fails to produce. RosettaSearch's designs show improvements in structural fidelity metrics ranging between 18% to 68%, translating to a 2.5x improvement in design success rate. We observe that these gains in success rate are robust when RosettaSearch-designed sequences are evaluated with an independent structure prediction oracle (Chai-1) and generalize across two distinct LLM families (o4-mini and Gemini-3), with performance scaling consistently with reasoning capability. We further demonstrate that RosettaSearch improves the sequence fidelity of ProteinMPNN designs for de novo backbones from the Dayhoff atlas, showing that the approach generalizes beyond native protein structures to computationally generated backbones. We also demonstrate a multi-modal extension of RosettaSearch with vision-language models, where images of predicted protein structures are used as feedback to incorporate structural context to guide protein sequence generation. To our knowledge, this is the first large-scale demonstration that LLMs can serve as effective generative optimizers for backbone-conditioned protein sequence design, yielding systematic gains without any model retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。