arXiv:2508.02137cs.LGcs.AI2025-08

AuroBind通过结构-功能联合建模,实现百万级化合物超快筛选。

Fitness aligned structural modeling enables scalable virtual screening with AuroBind

  • 用百万级数据微调原子级模型,结合偏好优化与师生加速策略。
  • 在10个靶点上实测命中率7%-69%,部分化合物达皮摩尔级活性。
  • 适合药物研发人员快速筛选新药候选,尤其针对难成药靶点。

绝大多数人类蛋白尚未被药物靶向,超过96%的蛋白未被现有疗法覆盖。基于结构的虚拟筛选本可拓展可成药蛋白组,但现有方法缺乏原子级精度且无法预测结合适配度,限制了转化应用。我们提出AuroBind,一种可扩展的虚拟筛选框架,通过在百万级化药基因组数据上微调定制的原子级结构模型。AuroBind融合直接偏好优化、高置信度复合物自蒸馏及师生加速策略,联合预测配体结合结构与结合适配度。所提模型在结构与功能基准上均超越现有最优方法,实现对超大规模化合物库的10万倍加速筛选。在10个疾病相关靶点的前瞻性筛选中,命中率达7%-69%,最优化合物表现出亚纳摩尔至皮摩尔级效力。对于孤儿G蛋白偶联受体GPR151和GPR160,AuroBind成功识别激动剂与拮抗剂,成功率16%-30%,功能实验验证了其在肝癌与前列腺癌模型中的调控作用。AuroBind为结构-功能学习与高通量分子筛选提供通用框架,弥合结构预测与治疗发现间的鸿沟。

原文摘要 · Abstract (English)

Most human proteins remain undrugged, over 96% of human proteins remain unexploited by approved therapeutics. While structure-based virtual screening promises to expand the druggable proteome, existing methods lack atomic-level precision and fail to predict binding fitness, limiting translational impact. We present AuroBind, a scalable virtual screening framework that fine-tunes a custom atomic-level structural model on million-scale chemogenomic data. AuroBind integrates direct preference optimization, self-distillation from high-confidence complexes, and a teacher-student acceleration strategy to jointly predict ligand-bound structures and binding fitness. The proposed models outperform state-of-the-art models on structural and functional benchmarks while enabling 100,000-fold faster screening across ultra-large compound libraries. In a prospective screen across ten disease-relevant targets, AuroBind achieved experimental hit rates of 7-69%, with top compounds reaching sub-nanomolar to picomolar potency. For the orphan GPCRs GPR151 and GPR160, AuroBind identified both agonists and antagonists with success rates of 16-30%, and functional assays confirmed GPR160 modulation in liver and prostate cancer models. AuroBind offers a generalizable framework for structure-function learning and high-throughput molecular screening, bridging the gap between structure prediction and therapeutic discovery.

虚拟筛选结构建模药物发现G蛋白偶联受体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。