arXiv:2409.03773q-bio.BMcs.LG2024-09AAAI被引 9

用跨域预训练模型提升蛋白质-核酸结合亲和力预测精度

CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction

  • 构建双模态融合架构,整合序列与结构信息
  • 在最大规模数据集PRA310上达到顶尖性能
  • 适合生物药物设计与突变效应分析研究者

精确测量蛋白质-核酸结合亲和力对生命过程研究和药物设计至关重要。以往方法仅依赖序列或结构特征,难以全面捕捉结合机制。近期基于海量蛋白质与核酸序列的预训练语言模型在同域任务中表现优异,但跨域模型协同用于复杂任务仍属空白。本文提出CoPRA,通过复杂结构连接不同生物域预训练模型,首次实现跨模态语言模型协作提升结合亲和力预测能力。提出Co-Former融合跨模态序列与结构信息,并设计双尺度预训练策略增强交互理解。构建了目前最大的蛋白质-核酸结合亲和力数据集PRA310用于评估。在公开突变效应预测数据集上亦验证有效性。CoPRA在所有数据集上均达领先水平。大量分析表明其可准确预测结合亲和力、解析突变引起的亲和力变化,并随数据量与模型规模扩大持续受益。

原文摘要 · Abstract (English)

Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent emerging pre-trained language models trained on massive unsupervised sequences of protein and RNA have shown strong representation ability for various in-domain downstream tasks, including binding site prediction. However, applying different-domain language models collaboratively for complex-level tasks remains unexplored. In this paper, we propose CoPRA to bridge pre-trained language models from different biological domains via Complex structure for Protein-RNA binding Affinity prediction. We demonstrate for the first time that cross-biological modal language models can collaborate to improve binding affinity prediction. We propose a Co-Former to combine the cross-modal sequence and structure information and a bi-scope pre-training strategy for improving Co-Former's interaction understanding. Meanwhile, we build the largest protein-RNA binding affinity dataset PRA310 for performance evaluation. We also test our model on a public dataset for mutation effect prediction. CoPRA reaches state-of-the-art performance on all the datasets. We provide extensive analyses and verify that CoPRA can (1) accurately predict the protein-RNA binding affinity; (2) understand the binding affinity change caused by mutations; and (3) benefit from scaling data and model size.

蛋白质-核酸结合预训练模型亲和力预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。