arXiv:2505.12913cs.LGq-bio.QM2025-05被引 3

用片段主动学习高效探索万亿级分子空间,提升药物设计效率。

Active Learning on Synthons for Molecular Design

  • 基于片段选择的主动学习框架,突破传统枚举限制。
  • 仅用少量样本即可在万亿级分子空间中找到高分候选分子。
  • 适合多目标药物分子设计,优于主流生成方法且更富多样性。

全面虚拟筛选虽信息丰富,但在现代药物发现中因计算成本过高而难以实施。这一问题在多向量扩展等组合场景下尤为严重,导致分子空间迅速膨胀至超大规模。本文提出可扩展的片段获取主动学习方法(SALSA):一种适用于多向量扩展的简单算法,通过将建模与采样过程分解为片段(synthon)选择,将池基主动学习推广至非枚举空间。在配体和结构基础目标上的实验表明,SALSA具有出色的样本效率,并能扩展至包含万亿级化合物的空间。进一步在三个蛋白靶点上验证了其在多参数目标设计中的应用,结果表明SALSA生成分子的化学性质与已知活性分子相当,且多样性更高、得分优于行业领先生成方法。

原文摘要 · Abstract (English)

Exhaustive virtual screening is highly informative but often intractable against the expensive objective functions involved in modern drug discovery. This problem is exacerbated in combinatorial contexts such as multi-vector expansion, where molecular spaces can quickly become ultra-large. Here, we introduce Scalable Active Learning via Synthon Acquisition (SALSA): a simple algorithm applicable to multi-vector expansion which extends pool-based active learning to non-enumerable spaces by factoring modeling and acquisition over synthon or fragment choices. Through experiments on ligand- and structure-based objectives, we highlight SALSA's sample efficiency, and its ability to scale to spaces of trillions of compounds. Further, we demonstrate application toward multi-parameter objective design tasks on three protein targets - finding SALSA-generated molecules have comparable chemical property profiles to known bioactives, and exhibit greater diversity and higher scores over an industry-leading generative approach.

分子设计主动学习生成模型药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。