用预结构嵌入预测蛋白-RNA结合亲和力,无需三维结构
ZeroFold: Protein-RNA Binding Affinity Predictions from Pre-Structural Embeddings
- 利用生物分子基础模型的预结构嵌入捕捉RNA构象动态
- 在2621对数据上达0.65的斯皮尔曼相关系数,接近实验噪声上限
- 适合无结构信息的蛋白-RNA配对亲和力预测,尤其低相似性场景
蛋白质-RNA结合亲和力的精确预测仍是结构生物学中的未解难题,限制了对基因调控的理解和靶向RNA药物的设计。主要障碍在于RNA的结构柔性——与蛋白质不同,RNA以动态构象集合形式存在。若仅依赖单一预测结构,则会丢失关键结合信息。本文提出通过提取预结构嵌入(pre-structural embeddings)解决该问题:这些中间表示来自生物分子基础模型,在结构解码前捕获,隐式编码构象集合信息而无需实际结构。我们构建了ZeroFold,一个基于Transformer的模型,通过跨模态注意力机制融合Boltz-2生成的蛋白与RNA预结构嵌入,直接从序列预测结合亲和力。为支持训练与评估,我们构建了PRADB数据集,包含4个互补数据库中收集的2,621对独特蛋白-RNA复合物,其结合亲和力均有实验测定。在保留40%序列同一性阈值的测试集上,ZeroFold取得0.65的斯皮尔曼相关系数,接近由实验测量噪声决定的上限。在逐步更公平的评估条件下(控制训练集重叠),相较于领先的基于结构和基于序列的预测器,ZeroFold表现更优,且随着与竞争模型训练数据的序列相似度降低,性能差距持续扩大。结果表明,预结构嵌入为柔性生物分子提供了一种有效表征策略,为无结构数据的蛋白-RNA配对亲和力预测开辟新路径。
原文摘要 · Abstract (English)
The accurate prediction of protein-RNA binding affinity remains an unsolved problem in structural biology, limiting opportunities in understanding gene regulation and designing RNA-targeting therapeutics. A central obstacle is the structural flexibility of RNA, as, unlike proteins, RNA molecules exist as dynamic conformational ensembles. Thus, committing to a single predicted structure discards information relevant to binding. Here, we show that this obstacle can be addressed by extracting pre-structural embeddings, which are intermediate representations from a biomolecular foundation model captured before the structure decoding step. Pre-structural embeddings implicitly encode conformational ensemble information without requiring predicted structures. We build ZeroFold, a transformer-based model that combines pre-structural embeddings from Boltz-2 for both protein and RNA molecules through a cross-modal attention mechanism to predict binding affinity directly from sequence. To support training and evaluation, we construct PRADB, a curated dataset of 2,621 unique protein-RNA pairs with experimentally measured affinities drawn from four complementary databases. On a held-out test set constructed with 40% sequence identity thresholds, ZeroFold achieves a Spearman correlation of 0.65, a value approaching the ceiling imposed by experimental measurement noise. Under progressively fairer evaluation conditions that control for training-set overlap, ZeroFold compares favourably with respect to leading structure-based and leading sequence-based predictors, with the performance gap widening as sequence similarity to competitor training data is reduced. These results illustrate how pre-structural embeddings offer a representation strategy for flexible biomolecules, opening a route to affinity prediction for protein-RNA pairs for which no structural data exist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。