用分子嵌入选对接种算法,提升蛋白-配体对接效果
Molecular Embedding-Based Algorithm Selection in Protein-Ligand Docking
- 基于蛋白与配体嵌入,用注意力池化和浅残差解码预测算法表现
- 比最优单算法提升15个百分点,缩小虚拟最优与实际最优差距17%~66%
- 适合固定流程场景,可诊断算法选择是否合理
选择有效的对接算法高度依赖具体情境,单一方法在不同结构、化学和协议条件下表现不稳定。MolAS 是一个轻量级算法选择模型,通过预训练的蛋白和配体嵌入,利用注意力池化和浅残差解码器,预测各算法的性能。在数百至数千个标注复合物上,MolAS 相比单个最优求解器(SBS)实现高达15个百分点的绝对性能提升,并在五个对接基准上缩小了虚拟最优求解器(VBS)与SBS之间的差距17%~66%。对选择频率、条件边际可靠性及基准级预言机结构的分析表明,当工作流定义的预言机景观具有低胜者熵且顶层求解器区域可区分时,MolAS 效果最佳;但在协议不匹配导致求解器排名变化时性能下降。结果表明,在评估范围内,鲁棒性受限于工作流与协议引起的求解器层级不稳定,而非表征能力不足,因此 MolAS 适合作为固定流程中的域内选择器,或用于评估对接算法选择问题是否成立。
原文摘要 · Abstract (English)
Selecting an effective docking algorithm is highly context-dependent, and no single method performs reliably across structural, chemical, and protocol regimes. MolAS is a lightweight algorithm-selection model that predicts per-algorithm performance from pretrained protein and ligand embeddings using attentional pooling and a shallow residual decoder. With hundreds to a few thousand labelled complexes, MolAS achieves up to a 15 percentage-point absolute improvement over the single-best solver (SBS) and closes 17--66\% of the Virtual Best Solver (VBS)--SBS gap across five docking benchmarks. Analyses of selection frequencies, margin-conditioned reliability, and benchmark-level oracle structure indicate that MolAS is most effective when the workflow-defined oracle landscape has low winner entropy and a reasonably separable top-solver region, but degrades under protocol mismatch that shifts solver rankings and changes the induced labels. These results suggest that, in the evaluated regime, robustness is limited less by representational capacity than by workflow- and protocol-induced instability in solver hierarchies, positioning MolAS as an in-domain selector for fixed pipelines and as a diagnostic tool for assessing when docking algorithm selection is well-posed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。