CalPro让蛋白质结构预测更可靠,抗分布偏移,误差更小。
CalPro: Prior-Aware Evidential--Conformal Prediction with Structure-Aware Guarantees for Protein Structures
- 用图结构输出概率分布,结合领域先验做约束
- 跨模态覆盖误差仅降5%(基线15-25%),校准误差降30-50%
- 适合需要高可信度结构预测的药物设计等场景
深度蛋白质结构预测模型如AlphaFold提供的置信度估计(如pLDDT)常存在校准偏差,且在实验模态、时间变化及内在无序区域下性能下降。我们提出CalPro,一种先验感知的证据-共形预测框架,实现对分布偏移的鲁棒不确定性量化。CalPro结合:(i) 基于图结构的几何证据头,输出正态逆伽马预测分布;(ii) 可微分共形层,支持端到端训练并保证有限样本覆盖;(iii) 将无序性、柔韧性等域先验作为软约束编码。通过在模糊集上使用PAC-Bayesian界,推导出结构感知的覆盖保证。实证表明,CalPro在跨模态下覆盖损失不超过5%(基线15-25%),校准误差降低30-50%,下游配体对接成功率提升25%。该方法亦适用于其他具有局部可靠性先验的结构化回归任务,在非生物基准上验证有效。
原文摘要 · Abstract (English)
Deep protein structure predictors such as AlphaFold provide confidence estimates (e.g., pLDDT) that are often miscalibrated and degrade under distribution shifts across experimental modalities, temporal changes, and intrinsically disordered regions. We introduce CalPro, a prior-aware evidential-conformal framework for shift-robust uncertainty quantification. CalPro combines (i) a geometric evidential head that outputs Normal-Inverse-Gamma predictive distributions via a graph-based architecture; (ii) a differentiable conformal layer that enables end-to-end training with finite-sample coverage guarantees; and (iii) domain priors (disorder, flexibility) encoded as soft constraints. We derive structure-aware coverage guarantees under distribution shift using PAC-Bayesian bounds over ambiguity sets, and show that CalPro maintains near-nominal coverage while producing tighter intervals than standard conformal methods in regions where priors are informative. Empirically, CalPro exhibits at most 5% coverage degradation across modalities (vs. 15-25% for baselines), reduces calibration error by 30-50%, and improves downstream ligand-docking success by 25%. Beyond proteins, CalPro applies to structured regression tasks in which priors encode local reliability, validated on non-biological benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。