arXiv:2603.14350cs.LG2026-03KDD

融合结构数据库与深度学习,提升蛋白质逆折叠设计精度

Refold: Refining Protein Inverse Folding with Efficient Structural Matching and Fusion

  • 结合数据库结构先验与模型预测,动态融合优化残基概率
  • 在CATH 4.2和4.3上实现0.63的天然序列恢复率,达到当前最优
  • 特别提升高不确定性区域性能,适合复杂蛋白设计任务

蛋白质逆折叠旨在设计能折叠成给定主链结构的氨基酸序列,是蛋白质设计的核心任务。现有方法分为两类:基于模板的方法依赖数据库结构先验,在邻近结构存在时局部精度高,但受数据库覆盖度和匹配质量限制,对分布外(OOD)目标表现差;深度学习方法虽泛化性强,却难以捕捉细微局部结构,导致残基预测不确定、遗漏局部基序。本文提出Refold框架,通过匹配邻居获取结构先验,并与模型预测融合以精炼残基概率。针对低质量先验引入噪声的问题,设计动态效用门控机制,当先验不可信时自动回退至基础预测。在标准基准上的全面评估显示,Refold在CATH 4.2和CATH 4.3上均实现0.63的天然序列恢复率,达到当前最优。分析表明,其在高不确定性区域提升更显著,体现了结构先验与深度学习互补的优势。

原文摘要 · Abstract (English)

Protein inverse folding aims to design an amino acid sequence that will fold into a given backbone structure, serving as a central task in protein design. Two main paradigms have been widely explored. Template-based methods exploit database-derived structural priors and can achieve high local precision when close structural neighbors are available, but their dependence on database coverage and match quality often degrades performance on out-of-distribution (OOD) targets. Deep learning approaches, in contrast, learn general structure-to-sequence regularities and usually generalize better to new backbones. However, they struggle to capture fine-grained local structure, which can cause uncertain residue predictions and missed local motifs in ambiguous regions. We introduce Refold, a novel framework that synergistically integrates the strengths of database-derived structural priors and deep learning prediction to enhance inverse folding. Refold obtains structural priors from matched neighbors and fuses them with model predictions to refine residue probabilities. In practice, low-quality neighbors can introduce noise, potentially degrading model performance. We address this issue with a Dynamic Utility Gate that controls prior injection and falls back to the base prediction when the priors are untrustworthy. Comprehensive evaluations on standard benchmarks demonstrate that Refold achieves state-of-the-art native sequence recovery of 0.63 on both CATH 4.2 and CATH 4.3. Also, analysis indicates that Refold delivers larger gains on high-uncertainty regions, reflecting the complementarity between structural priors and deep learning predictions.

蛋白质设计逆折叠结构融合深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。