arXiv:2605.28287cs.LGcond-mat.mtrl-sci2026-05

用强化学习从头探索化学空间,无需预训练数据即可发现新分子。

AtomComposer: Discovering Chemical Space from First Principles with Reinforcement Learning

论文配图:AtomComposer: Discovering Chemical Space from First Principles with Reinforcement Learning
图 1 · 摘自论文原文
  • 自导向代理通过在线强化学习构建三维异构体,不依赖预训练数据。
  • 在未见分子式上生成的有效异构体数量比基线多一个数量级。
  • 适用于新药研发、材料设计等需要突破传统数据限制的领域。

无训练数据条件下发现新稳定分子仍是重大科学挑战。现有分子生成模型依赖大规模预标注数据集,引入偏见并限制新化学领域的探索。与此相反,我们提出一种新范式:无需预训练的自主通用智能体,可独立探索广阔未知化学空间。首次提出AtomComposer,一种自引导代理,能在线仅通过强化学习,在化学计量约束下自主构建有效3D异构体。不同于通常对特定化学式过拟合的方法,我们采用多组成训练策略,以能量和有效性奖励为引导,实现跨多样化化学的广泛泛化。在未见测试公式上,该代理发现的有效异构体数量比现有单组成强化学习基线高出一个数量级。这些结果验证了在线强化学习作为可扩展、从零开始探索化学构型空间的强大范式。

原文摘要 · Abstract (English)

Discovering novel stable molecules without training data remains a grand scientific challenge. Current molecular generative models are trained on large, pre-curated datasets, which introduce biases and limit exploration of novel chemistry. In contrast, we propose a new paradigm: autonomous, generalized agents capable of mapping vast, unknown chemical spaces without any pretraining. For the first time, we present AtomComposer, a self-guided agent that autonomously constructs valid 3D isomers under stoichiometric constraints and is trained exclusively online using reinforcement learning. Unlike existing approaches that generally overfit to a specific chemical formula, we establish a multi-composition training scheme that enables a broad generalization across diverse chemistry, guided by energy- and validity-based rewards. Our agent can discover up to an order of magnitude more valid isomers on unseen test formulas than existing single-composition reinforcement-learning baselines trained with per-step energy rewards. These results fulfill the promise of online reinforcement learning as a powerful paradigm for scalable, from-scratch exploration of chemical configuration space.

分子生成强化学习化学空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。