arXiv:2506.09496cs.LGcs.AI2025-06被引 2

用能量引导设计更稳定的蛋白质序列,比现有方法更精准预测结构能量。

EnerBridge-DPO: Energy-Guided Protein Inverse Folding with Markov Bridges and Direct Preference Optimization

  • 结合马尔可夫桥与偏好优化,从高信息量序列出发生成候选
  • 引入能量约束损失,直接学习并预测序列能量值
  • 在保持高恢复率的同时显著降低能量,适合蛋白设计研究者

蛋白质逆折叠中设计具有最优能量稳定性的序列是关键挑战,当前深度学习方法主要以最大化序列恢复率为训练目标,常忽略生成序列的能量。本文提出EnerBridge-DPO框架,旨在直接生成低能量、高稳定性的蛋白质序列。核心创新包括:首先,将马尔可夫桥与直接偏好优化(DPO)结合,利用基于能量的偏好对马尔可夫桥模型进行微调,马尔可夫桥从富含信息的先验序列启动优化,为DPO提供结构合理的序列候选池;其次,引入显式能量约束损失,强化基于先验序列的DPO能量驱动特性,使模型能有效从大量先验知识中学习能量表征,并直接预测序列能量值,从而捕捉能量景观的定量特征。评估表明,EnerBridge-DPO在生成复杂蛋白质序列时能量更低,同时保持与最先进模型相当的序列恢复率,并准确预测不同序列间的$ΔΔG$值。

原文摘要 · Abstract (English)

Designing protein sequences with optimal energetic stability is a key challenge in protein inverse folding, as current deep learning methods are primarily trained by maximizing sequence recovery rates, often neglecting the energy of the generated sequences. This work aims to overcome this limitation by developing a model that directly generates low-energy, stable protein sequences. We propose EnerBridge-DPO, a novel inverse folding framework focused on generating low-energy, high-stability protein sequences. Our core innovation lies in: First, integrating Markov Bridges with Direct Preference Optimization (DPO), where energy-based preferences are used to fine-tune the Markov Bridge model. The Markov Bridge initiates optimization from an information-rich prior sequence, providing DPO with a pool of structurally plausible sequence candidates. Second, an explicit energy constraint loss is introduced, which enhances the energy-driven nature of DPO based on prior sequences, enabling the model to effectively learn energy representations from a wealth of prior knowledge and directly predict sequence energy values, thereby capturing quantitative features of the energy landscape. Our evaluations demonstrate that EnerBridge-DPO can design protein complex sequences with lower energy while maintaining sequence recovery rates comparable to state-of-the-art models, and accurately predicts $ΔΔG$ values between various sequences.

蛋白设计能量预测生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。