改进蛋白语言模型的架构,提升多链蛋白相互作用亲和力预测精度
Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction
- 提出四种新架构,避免简单拼接蛋白表征
- 最高使相关系数提升12%,显著优于传统方法
- 适合蛋白质互作研究与药物设计领域的研究人员
蛋白质-蛋白质相互作用(PPI)在众多细胞过程中起关键作用,其特征描述对理解疾病机制和指导药物发现至关重要。尽管蛋白质语言模型(PLMs)在预测蛋白质结构和功能方面表现优异,但其在基于序列的PPI结合亲和力预测中的应用仍相对不足。这一差距常归因于高质量、严格筛选数据集的缺乏以及依赖简单的蛋白表征拼接策略。本文针对这些问题:首先,我们构建了一个经仔细整理的PPB-Affinity数据集,共8,207个唯一多链蛋白相互作用条目,通过修正注释不一致和重复条目,并采用≤30%序列相似性阈值进行严格划分训练、验证和测试集,有效防止数据泄露;其次,系统评估了四种适配PLMs用于PPI亲和力预测的架构:嵌入拼接(EC)、序列拼接(SC)、分层池化(HP)和池化注意力相加(PAD)。实验涵盖两种训练方式:全微调和仅在冻结的PLM特征上使用ConvBERT头的轻量级方法。在五种主流PLMs(ProtT5、ESM2、Ankh、Ankh2、ESM3)上的全面实验表明,HP和PAD架构持续优于传统拼接方法,在斯皮尔曼相关系数上最高提升达12%,凸显了复杂架构设计对充分挖掘PLMs在精细PPI亲和力预测中潜力的重要性。
原文摘要 · Abstract (English)
Protein-protein interactions (PPIs) are fundamental to numerous cellular processes, and their characterization is vital for understanding disease mechanisms and guiding drug discovery. While protein language models (PLMs) have demonstrated remarkable success in predicting protein structure and function, their application to sequence-based PPI binding affinity prediction remains relatively underexplored. This gap is often attributed to the scarcity of high-quality, rigorously refined datasets and the reliance on simple strategies for concatenating protein representations. In this work, we address these limitations. First, we introduce a meticulously curated version of the PPB-Affinity dataset of a total of 8,207 unique protein-protein interaction entries, by resolving annotation inconsistencies and duplicate entries for multi-chain protein interactions. This dataset incorporates a stringent, less than or equal to 30%, sequence identity threshold to ensure robust splitting into training, validation, and test sets, minimizing data leakage. Second, we propose and systematically evaluate four architectures for adapting PLMs to PPI binding affinity prediction: embeddings concatenation (EC), sequences concatenation (SC), hierarchical pooling (HP), and pooled attention addition (PAD). These architectures were assessed using two training methods: full fine-tuning and a lightweight approach employing ConvBERT heads over frozen PLM features. Our comprehensive experiments across multiple leading PLMs (ProtT5, ESM2, Ankh, Ankh2, and ESM3) demonstrated that the HP and PAD architectures consistently outperform conventional concatenation methods, achieving up to 12% increase in terms of Spearman correlation. These results highlight the necessity of sophisticated architectural designs to fully exploit the capabilities of PLMs for nuanced PPI binding affinity prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。