arXiv:2509.00704cs.LGcs.AI2025-09中稿 · NeurIPS被引 1

用生成模型替代传统采样,让药物筛选更高效

Why Pool When You Can Flow? Active Learning with GFlowNets

  • 用GFlowNets直接按重要性采样分子,跳过全量评估
  • 在十亿级分子库中表现媲美传统方法,多样性更高
  • 适合大规模虚拟筛选,尤其药企做新药发现

池式主动学习的可扩展性受限于对大规模未标记数据集的计算评估成本,这一问题在药物发现的虚拟筛选中尤为突出。尽管贝叶斯主动学习通过分歧(BALD)等策略优先选择信息量大的样本,但在包含数十亿分子的数据库上仍计算开销巨大。本文提出BALD-GFlowNet,一种基于生成流网络(GFlowNets)的主动学习框架,可直接按BALD奖励比例采样目标对象。通过以生成采样取代传统池式选择,该方法实现与未标记数据池规模无关的可扩展性。在虚拟筛选实验中,BALD-GFlowNet性能与标准BALD相当,同时生成更具结构多样性的分子,为高效、可扩展的分子发现提供了新路径。

原文摘要 · Abstract (English)

The scalability of pool-based active learning is limited by the computational cost of evaluating large unlabeled datasets, a challenge that is particularly acute in virtual screening for drug discovery. While active learning strategies such as Bayesian Active Learning by Disagreement (BALD) prioritize informative samples, it remains computationally intensive when scaled to libraries containing billions samples. In this work, we introduce BALD-GFlowNet, a generative active learning framework that circumvents this issue. Our method leverages Generative Flow Networks (GFlowNets) to directly sample objects in proportion to the BALD reward. By replacing traditional pool-based acquisition with generative sampling, BALD-GFlowNet achieves scalability that is independent of the size of the unlabeled pool. In our virtual screening experiment, we show that BALD-GFlowNet achieves a performance comparable to that of standard BALD baseline while generating more structurally diverse molecules, offering a promising direction for efficient and scalable molecular discovery.

主动学习生成模型药物发现GFlowNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。