新评估方法可发现蛋白折叠中的隐蔽错误,提升结构预测可靠性。
CONFIDE: Hallucination Assessment for Reliable Biomolecular Structure Prediction and Design
- 用扩散嵌入自监督分析折叠拓扑困境,不依赖标注数据。
- 相比pLDDT,对折叠速率的预测相关性提升148%(0.82 vs 0.33)。
- 适用于药物设计、突变效应预测等任务,适合结构生物学研究者。
可靠评估蛋白质结构预测仍具挑战,因pLDDT虽反映能量稳定性,却常忽略原子碰撞或构象陷阱等细微错误,这些错误体现于蛋白质折叠能量景观中的拓扑挫折。我们提出CODE(Chain of Diffusion Embeddings),一种从AlphaFold3系列预测器的潜在扩散嵌入中,以全无监督方式直接量化拓扑挫折的自评估指标。结合pLDDT,我们构建了统一评估框架CONFIDE,融合能量与拓扑双重视角,提升AlphaFold3及相关模型的可靠性。CODE与由拓扑挫折驱动的折叠速率相关性达0.82,显著优于pLDDT的0.33(相对提升148%)。CONFIDE在分子胶结构预测基准中显著改善质量评估可靠性,与RMSD的斯皮尔曼相关性达0.73,相较pLDDT的0.42(相对提升73.8%)。该方法还可应用于全原子结合剂设计、酶活性位点定位、突变诱导结合亲和力预测、核酸适体筛选及柔性蛋白建模等多种药物设计任务。通过融合数据驱动嵌入与理论洞察,CODE与CONFIDE在多种生物分子系统中超越现有指标,为结构预测优化、结构生物学推进和药物发现加速提供稳健且通用的工具。
原文摘要 · Abstract (English)
Reliable evaluation of protein structure predictions remains challenging, as metrics like pLDDT capture energetic stability but often miss subtle errors such as atomic clashes or conformational traps reflecting topological frustration within the protein folding energy landscape. We present CODE (Chain of Diffusion Embeddings), a self evaluating metric empirically found to quantify topological frustration directly from the latent diffusion embeddings of the AlphaFold3 series of structure predictors in a fully unsupervised manner. Integrating this with pLDDT, we propose CONFIDE, a unified evaluation framework that combines energetic and topological perspectives to improve the reliability of AlphaFold3 and related models. CODE strongly correlates with protein folding rates driven by topological frustration, achieving a correlation of 0.82 compared to pLDDT's 0.33 (a relative improvement of 148\%). CONFIDE significantly enhances the reliability of quality evaluation in molecular glue structure prediction benchmarks, achieving a Spearman correlation of 0.73 with RMSD, compared to pLDDT's correlation of 0.42, a relative improvement of 73.8\%. Beyond quality assessment, our approach applies to diverse drug design tasks, including all-atom binder design, enzymatic active site mapping, mutation induced binding affinity prediction, nucleic acid aptamer screening, and flexible protein modeling. By combining data driven embeddings with theoretical insight, CODE and CONFIDE outperform existing metrics across a wide range of biomolecular systems, offering robust and versatile tools to refine structure predictions, advance structural biology, and accelerate drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。