arXiv:2602.18476q-bio.BMcs.AI2026-02

用语言模型提升蛋白-配体结合打分的精度与泛化能力

BioLM-Score: Language-Prior Conditioned Probabilistic Geometric Potentials for Protein-Ligand Scoring

  • 融合语言模型增强结构表征,联合建模蛋白与配体几何关系
  • 在CASF-2016上实现打分、排序、筛选等任务显著提升
  • 兼具高效性、跨靶点泛化性与可解释性,适合药物发现场景

蛋白-配体打分是基于结构的药物设计核心,支撑分子对接、虚拟筛选和构象优化。传统物理能量函数计算成本高,深度学习模型虽效率高但泛化能力差且可解释性弱。本文提出BioLM-Score,一种将几何建模与表示学习结合的通用打分模型。该模型采用针对蛋白与配体的模态专用、结构感知编码器,并引入生物分子语言模型丰富结构与化学表征。随后通过混合密度网络整合表征,预测多模态原子间距离分布,进而生成基于统计似然的打分。在CASF-2016基准测试中,BioLM-Score在对接、打分、排序与筛选任务上均取得显著提升。此外,该打分函数可有效作为引导对接流程与构象搜索的优化目标。总体而言,BioLM-Score为结构药物发现提供了一种兼具效率、泛化性与可解释性的实用替代方案。

原文摘要 · Abstract (English)

Protein-ligand scoring is a central component of structure-based drug design, underpinning molecular docking, virtual screening, and pose optimization. Conventional physics-based energy functions are often computationally expensive, limiting their utility in large-scale screening. In contrast, deep learning-based scoring models offer improved computational efficiency but frequently suffer from limited cross-target generalization and poor interpretability, which restrict their practical applicability. Here we present BioLM-Score, a simple yet generalizable protein-ligand scoring model that couples geometric modeling with representation learning. Specifically, it employs modality-specific and structure-aware encoders for proteins and ligands, each augmented with biomolecular language models to enrich structural and chemical representations. Subsequently, these representations are integrated through a mixture density network to predict multimodal interatomic distance distributions, from which statistically grounded likelihood-based scores are derived. Evaluations on the CASF-2016 benchmark demonstrate that BioLM-Score achieves significant improvements across docking, scoring, ranking, and screening tasks. Moreover, the proposed scoring function serves as an effective optimization objective for guiding docking protocols and conformational search. In summary, BioLM-Score provides a principled and practical alternative to existing scoring functions, combining efficiency, generalization, and interpretability for structure-based drug discovery.

蛋白打分药物设计语言模型几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。