arXiv:2605.24045cs.LGcs.AI2026-05

新数据集揭示蛋白-配体模型更懂结合概率而非结合位点。

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

论文配图:A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?
图 1 · 摘自论文原文
  • 构建10万对蛋白-配体数据,评估模型定位结合位点能力
  • 模型在结合预测上表现强,但定位结合位点能力普遍不足
  • 适合关注可解释性与物理机制的药物设计研究者

蛋白-配体建模是计算药物发现和分子设计的基础。现有基准多聚焦于判断蛋白与配体是否结合及结合强度,如二分类结合预测和亲和力回归任务,但难以检验模型是否能准确定位结合位点或识别非共价相互作用机制。为此,我们提出InteractBind,一个包含约10万对蛋白-配体的数据集,并配套细粒度评估基准。核心任务为结合位点定位,通过六类主要非共价相互作用的残基-原子作用图,评估模型生成的作用图能否准确聚焦结合区域。InteractBind还包含结合亲和力和蛋白相似性控制的划分,支持真实泛化能力评估。基于该数据集,我们评估了八种基于序列和交互感知的模型,涵盖二分类结合预测与结合位点定位任务。结果表明:尽管模型在二分类任务中表现良好,但在结合位点定位上能力有限,且不同非共价相互作用类型间差异显著。总体而言,InteractBind确立了一种鼓励开发更具可解释性和物理基础的蛋白-配体模型的新基准范式。

原文摘要 · Abstract (English)

Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a protein and ligand interact and how strongly they bind, through tasks such as binary binding prediction and affinity regression. However, these evaluations provide limited evidence of whether models can localize binding sites or identify the non-covalent interactions underlying molecular recognition. To address this gap, we introduce InteractBind, a large-scale protein-ligand dataset comprising approximately 100k protein-ligand pairs, together with a benchmark for fine-grained evaluation. The core fine-grained task is that of binding-site localization, which uses protein-residue and ligand-atom interaction maps spanning six major types of non-covalent interactions to assess whether model-derived interaction maps localize binding sites. InteractBind further includes binding affinity and protein similarity-controlled splits to support realistic generalization assessment. Using InteractBind, we evaluate eight existing sequence-based and interaction-aware models, assessing binary binding prediction and binding-site localization. Results reveal limited binding-site localization despite strong binary binding prediction, with marked variation across non-covalent interaction types. Overall, InteractBind establishes a benchmark paradigm that encourages the development of more interpretable and physically grounded protein-ligand models.

蛋白-配体结合位点药物发现可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。