构建首个针对雌激素受体α的大规模高精度结合能数据集,推动机器学习药物设计发展。
ToxBench: A Binding Affinity Prediction Benchmark with AB-FEP-Calculated Labels for Human Estrogen Receptor Alpha
- 基于AB-FEP计算8770个复合物结合能,确保数据高精度
- 提出DualBind模型,在保持低误差(1.75 kcal/mol)前提下大幅降低计算成本
- 提供非重叠分组,适合评估模型泛化能力,适合药企和研究者使用
蛋白质-配体结合亲和力预测对药物研发与毒性评估至关重要。尽管机器学习(ML)有望实现快速准确的预测,但其进展受限于可靠数据的缺乏。而基于物理的方法如绝对结合自由能微扰法(AB-FEP)虽精度高,却因计算成本过高难以用于高通量场景。为此,我们推出ToxBench,首个专注于关键药靶人类雌激素受体α(ERα)的大规模AB-FEP数据集。该数据集包含8,770个ERα-配体复合物结构,结合自由能通过AB-FEP计算得出,其中子集经实验亲和力验证,均方根误差为1.75 kcal/mol;同时提供非重叠配体划分,以评估模型泛化能力。基于ToxBench,我们对主流机器学习方法进行基准测试,特别提出双损失框架的DualBind模型,有效学习结合能函数。结果表明,DualBind表现优异,且机器学习可在极低计算成本下逼近AB-FEP性能。
原文摘要 · Abstract (English)
Protein-ligand binding affinity prediction is essential for drug discovery and toxicity assessment. While machine learning (ML) promises fast and accurate predictions, its progress is constrained by the availability of reliable data. In contrast, physics-based methods such as absolute binding free energy perturbation (AB-FEP) deliver high accuracy but are computationally prohibitive for high-throughput applications. To bridge this gap, we introduce ToxBench, the first large-scale AB-FEP dataset designed for ML development and focused on a single pharmaceutically critical target, Human Estrogen Receptor Alpha (ER$α$). ToxBench contains 8,770 ER$α$-ligand complex structures with binding free energies computed via AB-FEP with a subset validated against experimental affinities at 1.75 kcal/mol RMSE, along with non-overlapping ligand splits to assess model generalizability. Using ToxBench, we further benchmark state-of-the-art ML methods, and notably, our proposed DualBind model, which employs a dual-loss framework to effectively learn the binding energy function. The benchmark results demonstrate the superior performance of DualBind and the potential of ML to approximate AB-FEP at a fraction of the computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。