arXiv:2606.24990cs.LGcs.AI2026-06

让化学生成模型更可靠,通过考虑预测不确定性避免盲目追求高分分子。

Uncertainty-aware reinforcement learning for chemical language models

论文配图:Uncertainty-aware reinforcement learning for chemical language models
图 1 · 摘自论文原文
  • 引入预测不确定性作为优化目标,平衡探索与可靠性
  • 使真实命中率从0.5提升至0.75,真阳性数几乎翻倍
  • 适用于需要稳定生成的药物设计场景

强化学习(RL)已成为从头分子设计的强大工具,使化学语言模型(CLMs)能够探索化学空间并优化特定属性。然而,现有框架将所有评分函数视为确定性代理,忽略了不同分子属性预测的固有不确定性。这可能导致模型在高不确定性区域探索,生成看似高分但数据支持不足的分子,从而干扰优化过程,产生远离真实值的预测。本文提出两种互补方法:一是将不确定性作为额外优化目标,与其它评分函数共同作用,使策略在利用与可靠性间权衡;二是用不确定性调节策略更新,降低远离评分函数置信域的分子的影响。在三种设置下评估:(i)模拟系统,预测误差服从高斯分布,方差与距离训练数据远近成正比;(ii)使用ChemProp模型;(iii)基于随机森林分类器的共形预测包装器。结果表明,不确定性感知的强化学习使模型更稳健地探索化学空间,优先选择低不确定性区域,提升真实命中率0.25(从0.5增至0.75),真阳性数量接近翻倍。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has become a powerful paradigm for de novo molecular design, enabling Chemical Language Models (CLMs) to navigate and explore the chemical space while optimizing specific desired properties. However, the existing RL frameworks treat all scoring functions as deterministic oracles, neglecting the inherent uncertainty attached to the predictions of the different molecular properties. This can lead to the exploration of highly-uncertain regions of the chemical space, focusing on the generation of highly scored molecules which are poorly supported by the training data. This can destabilize the optimization process, yielding predictions that are far from their true values. We propose and compare two complementary ways of incorporating predictive uncertainty into RL. In the first one, uncertainty is treated as an additional optimization objective and incorporated along with the rest of the scoring functions, allowing the policy to trade off exploitation against reliability. Secondly, uncertainty is used to modulate policy updates, reducing the influence of molecules whose properties lie far outside the scoring function confidence domain. Both approaches were evaluated across three different settings: (i) a controlled model system, in which the prediction error is modeled as a Gaussian distribution, with a variance proportional to the distance to the training data; and two real-world tasks, making use of (ii) ChemProp models and (iii) a Conformal Prediction wrapper applied to a Random forest classifier. We show that uncertainty-aware RL enables CLMs to explore chemical space more robustly by favoring lower-uncertainty regions. This leads to more reliable hit discovery without compromising molecular score, increasing the true hit rate by 0.25 (from 0.5 to 0.75), and nearly doubling the total number of true hits.

强化学习化学生成不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。