评测AI药物发现工具Boltz-2,发现其虽快但精度不足,难替代物理模型。
On the Reliability of AI Methods in Drug Discovery: Evaluation of Boltz-2 for Structure and Binding Affinity Prediction
- 用联合共折叠方法预测蛋白-配体结构与亲和力
- 对两大数据集测试,结构误差大,能量相关性弱
- 适合快速初筛,不适用于精准候选药识别
尽管人工智能在药物发现领域备受关注,但至今尚无基于AI发现的药物获得监管批准。本文评估了该领域最新工具Boltz-2在蛋白质-配体结构及结合亲和力预测方面的表现。该模型采用联合共折叠策略,旨在兼顾AI效率与物理精度。研究使用两个大规模数据集:16,780个化合物针对3CLPro,21,702个化合物针对TNKS2。通过与传统对接方法比较结构,并以基于物理的ESMACS协议计算的结合自由能作为基准,评估其预测能力。结构分析显示显著的全局RMSD差异,表明Boltz-2预测多种蛋白构象与配体结合位点,而非单一收敛构象。能量评估在整体数据集上仅呈现弱至中等相关性。对前100名化合物的聚焦分析进一步显示,Boltz-2预测结果与细粒度ESMACS结合自由能之间无显著相关性,且配体结构呈现饱和差异。结果表明,尽管Boltz-2在初步筛选中具备显著速度优势,但缺乏足够能量分辨率用于先导化合物识别。因此,需结合物理方法保障AI模型的可靠性与精细化。
原文摘要 · Abstract (English)
Despite continuing hype about the role of AI in drug discovery, no "AI-discovered drugs" have so far received regulatory approval. Here we assess one of the latest AI based tools in this domain. The ability to rapidly predict protein-ligand structures and binding affinities is pivotal for accelerating drug discovery. Boltz-2, a recently developed biomolecular foundation model, aims to bridge the gap between AI efficiency and physics-based precision through a joint "co-folding" approach. In this study, we provide an extensive evaluation of Boltz-2 using two large-scale datasets: 16,780 compounds for 3CLPro and 21,702 compounds for TNKS2. We compare Boltz-2 predicted structures with traditional docking and binding affinities with binding free energies derived from the physics-based ESMACS protocol. Structural analysis reveals significant global RMSD variations, indicating that Boltz-2 predicts multiple protein conformations and ligand binding positions rather than a single converged pose. Energetic evaluations exhibit only weak to moderate correlations across the global datasets. Furthermore, a focused analysis of the top 100 compounds yields no significant correlation between the Boltz-2 predictions and the binding free energies from fine-grained ESMACS, alongside observed saturation difference in ligand structures. Our results show that while Boltz-2 offers substantial speed for initial screening, it lacks the energetic resolution required for lead identification. These findings highlight the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。