arXiv:2506.17007cs.LG2025-06被引 2

提出新算法提升分子生成多样性与质量,避免低质候选物

Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning

  • 用通用软算子统一多种正则化强化学习方法
  • 在合成与真实任务中生成更优且多样的分子候选
  • 适合需要高效筛选高潜力分子的研究者

科学发现中的主要瓶颈在于从指数级庞大的对象集合(如蛋白质或分子)中筛选出少数具备理想性质的候选。传统方法依赖专家知识,而近年研究采用由代理奖励函数引导的强化学习来实现筛选。通过各类熵正则化,这些方法旨在学习生成多样且高分候选的采样器。本文指出,现有方法在大规模搜索空间中易生成过于多样但次优的候选。为此,我们提出一种新型统一算子,将多种正则化强化学习算子整合为更优的峰值采样框架。其次,我们从鲁棒强化学习视角重新审视该过程:正则化可视为对代理函数中组合性不确定性(即候选的真实评估与代理评估存在差异)的鲁棒性。基于此分析,我们提出简单易用的新算法——轨迹广义温和最大(TGM),在合成与真实任务中均优于基线,生成质量更高、多样性更好的候选。代码已开源。

原文摘要 · Abstract (English)

A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with desirable properties. While this process can rely on expert knowledge, recent methods leverage reinforcement learning (RL) guided by a proxy reward function to enable this filtering. By employing various forms of entropy regularization, these methods aim to learn samplers that generate diverse candidates that are highly rated by the proxy function. In this work, we make two main contributions. First, we show that these methods are liable to generate overly diverse, suboptimal candidates in large search spaces. To address this issue, we introduce a novel unified operator that combines several regularized RL operators into a general framework that better targets peakier sampling distributions. Secondly, we offer a novel, robust RL perspective of this filtering process. The regularization can be interpreted as robustness to a compositional form of uncertainty in the proxy function (i.e., the true evaluation of a candidate differs from the proxy's evaluation). Our analysis leads us to a novel, easy-to-use algorithm we name trajectory general mellowmax (TGM): we show it identifies higher quality, diverse candidates than baselines in both synthetic and real-world tasks. Code: https://github.com/marcojira/tgm.

强化学习分子生成优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。