用随机图编码器和物理特征提升蛋白质-配体结合亲和力预测精度
RAVEN: Frozen Random Graph Reservoirs with Physics-Informed Interaction Fingerprints for Protein-Ligand Binding Affinity Prediction
- 用多个冻结的原子图编码器生成多视角结构表示
- 在PDBbind和CASF-2016数据集上达到领先性能
- 适合需要高鲁棒性的药物分子亲和力预测场景
从三维复合物结构定量预测蛋白质-配体结合亲和力是基于结构的计算化学与分子建模中的基础任务。可靠预测仍具挑战性,因可用结构-亲和力数据有限、实验异质性强、构象依赖性高,且对数据划分敏感。RAVEN(随机原子视图与集成神经储层)采用多个独立初始化且完全冻结的原子图编码器,生成无需端到端优化的多样化结构投影。这些投影与确定性的物理化学相互作用指纹结合,由异构监督读取器(包括神经网络和树模型)处理,输出通过基于验证的非负融合整合。随机储层扩大了结构特征覆盖范围,而显式物理化学描述符和异构读取器提供互补信息与不同归纳偏置。在基于GEMS相似性资源重构的相似性隔离的PDBbind 2020R1划分,以及受保护的CASF-2016子集上的评估显示强大预测性能。结果表明,冻结的多视角图表示、显式物理化学统计及异构模型融合构成一个稳健且灵活的蛋白质-配体结合亲和力预测框架。
原文摘要 · Abstract (English)
Quantitative estimation of protein-ligand binding affinity from three-dimensional complex structures is a fundamental task in structure-based computational chemistry and molecular modeling. Reliable prediction remains challenging because available structure-affinity data are limited, experimentally heterogeneous, conformation-dependent, and sensitive to dataset partitioning. RAVEN (Randomized Atomistic Views with Ensemble Neural Reservoirs) utilizes a multihead reservoir of independently initialized and fully frozen atomistic graph encoders to generate diverse structural projections without end-to-end optimization of the graph representation. These projections are integrated with a deterministic physicochemical interaction fingerprint and processed by heterogeneous supervised readers, including neural and tree-based regressors, whose outputs are combined through validation-based nonnegative fusion. The random reservoir expands structural feature coverage across independent encoder realizations, whereas the explicit physicochemical descriptors and heterogeneous readers contribute complementary information and distinct inductive biases. Evaluation on a similarity-isolated PDBbind 2020R1 split reconstructed using GEMS similarity resources, together with the protected CASF-2016 subset, demonstrated strong predictive performance. The results indicate that frozen multi-view graph representations, explicit physicochemical statistics, and heterogeneous model fusion provide a robust and flexible framework for protein-ligand binding-affinity prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。