用结构模型Boltz-2预测蛋白-蛋白结合亲和力,效果不如序列模型
On fine-tuning Boltz-2 for protein-protein affinity prediction
- 将Boltz-2改用于蛋白-蛋白亲和力回归,基于结构特征建模
- 在TCR3d和PPB-affinity数据集上,性能低于序列模型
- 结构与序列嵌入融合可互补提升,尤其增强弱序列模型
准确预测蛋白-蛋白结合亲和力对理解分子相互作用和设计药物至关重要。我们将先进的基于结构的蛋白-配体亲和力预测模型Boltz-2 adapted用于蛋白-蛋白亲和力回归,并在TCR3d和PPB-affinity两个数据集上评估。尽管结构精度高,Boltz-2-PPI在小规模和大规模数据场景下均表现逊于序列基线模型。将Boltz-2-PPI的嵌入与序列模型嵌入结合,带来互补性提升,尤其对较弱的序列模型效果显著,表明序列与结构模型学习到的信号不同。结果呼应了使用结构数据训练时存在的已知偏差,提示当前结构表示尚未优化用于高效亲和力预测。
原文摘要 · Abstract (English)
Accurate prediction of protein-protein binding affinity is vital for understanding molecular interactions and designing therapeutics. We adapt Boltz-2, a state-of-the-art structure-based protein-ligand affinity predictor, for protein-protein affinity regression and evaluate it on two datasets, TCR3d and PPB-affinity. Despite high structural accuracy, Boltz-2-PPI underperforms relative to sequence-based alternatives in both small- and larger-scale data regimes. Combining embeddings from Boltz-2-PPI with sequence-based embeddings yields complementary improvements, particularly for weaker sequence models, suggesting different signals are learned by sequence- and structure-based models. Our results echo known biases associated with training with structural data and suggest that current structure-based representations are not primed for performant affinity prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。