arXiv:2602.07735cs.LGq-bio.BM2026-02被引 2

用粗粒度结构表示,实现快速精准的药物结合亲和力预测

TerraBind: Fast and Accurate Binding Affinity Prediction through Coarse Structural Representations

  • 采用蛋白主链和配体重原子的粗粒度表示,跳过耗时的全原子生成
  • 在多个数据集上亲和力预测相关性比现有方法提升约20%
  • 适合大规模药物筛选场景,尤其适用于需快速评估候选分子的团队

我们提出TerraBind,一种用于蛋白质-配体结构与结合亲和力预测的基础模型,在推理速度上比当前最优方法快26倍,同时将亲和力预测准确率提升约20%。现有深度学习方法依赖昂贵的全原子扩散模型生成3D坐标,导致大规模化合物筛选计算不可行。本文挑战这一范式,提出关键假设:精确的小分子构象与亲和力预测无需全原子分辨率。TerraBind通过多模态架构(结合COATI-3分子编码与ESM-2蛋白嵌入),在口袋级粗粒度表示(仅保留蛋白Cβ原子和配体重原子)下学习丰富结构表征,并在无扩散优化模块中生成构象,同时在亲和力概率预测模块中进行打分。在结构预测基准(FoldBench、PoseBusters、Runs N' Poses)上,其配体构象精度与基于扩散的方法相当。关键的是,在公开基准(CASP16)和多样化的专有数据集(18种生化/细胞实验)上,其亲和力预测的皮尔逊相关性比Boltz-2提升约20%。该模块还提供校准良好的亲和力不确定性估计,填补了药物发现中可靠化合物优先排序的关键空白。此外,该模块支持持续学习框架与风险规避批量选择策略,在模拟药物发现周期中,所选分子的亲和力提升效果是贪婪方法的6倍。

原文摘要 · Abstract (English)

We present TerraBind, a foundation model for protein-ligand structure and binding affinity prediction that achieves 26-fold faster inference than state-of-the-art methods while improving affinity prediction accuracy by $\sim$20\%. Current deep learning approaches to structure-based drug design rely on expensive all-atom diffusion to generate 3D coordinates, creating inference bottlenecks that render large-scale compound screening computationally intractable. We challenge this paradigm with a critical hypothesis: full all-atom resolution is unnecessary for accurate small molecule pose and binding affinity prediction. TerraBind tests this hypothesis through a coarse pocket-level representation (protein C$_β$ atoms and ligand heavy atoms only) within a multimodal architecture combining COATI-3 molecular encodings and ESM-2 protein embeddings that learns rich structural representations, which are used in a diffusion-free optimization module for pose generation and a binding affinity likelihood prediction module. On structure prediction benchmarks (FoldBench, PoseBusters, Runs N' Poses), TerraBind matches diffusion-based baselines in ligand pose accuracy. Crucially, TerraBind outperforms Boltz-2 by $\sim$20\% in Pearson correlation for binding affinity prediction on both a public benchmark (CASP16) and a diverse proprietary dataset (18 biochemical/cell assays). We show that the affinity prediction module also provides well-calibrated affinity uncertainty estimates, addressing a critical gap in reliable compound prioritization for drug discovery. Furthermore, this module enables a continual learning framework and a hedged batch selection strategy that, in simulated drug discovery cycles, achieves 6$\times$ greater affinity improvement of selected molecules over greedy-based approaches.

药物设计结合亲和力结构预测高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。