arXiv:2601.10131cs.AIcs.MA2026-01被引 3

用分段智能体精准生成满足多属性约束的分子

M^4olGen: Multi-Agent, Multi-Stage Molecular Generation under Precise Multi-Property Constraints

  • 分两阶段:先检索生成候选原型,再通过强化学习精细优化
  • 在多个属性上同时逼近目标值,有效性与精确性均优于现有方法
  • 适合药物研发中需严格控制理化性质的分子设计场景

生成同时满足多个物理化学属性精确数值约束的分子极具挑战性。尽管大语言模型表达能力强,但在缺乏外部结构与反馈的情况下难以实现精确的多目标控制与数值推理。我们提出 extbf{M^4olGen},一种基于片段级别的检索增强型两阶段分子生成框架。第一阶段:原型生成——多智能体推理器执行基于检索的片段级编辑,生成接近可行区域的候选分子;第二阶段:基于强化学习的细粒度优化——采用组相对策略优化(GRPO)训练的片段级优化器,通过单步或多次跳跃式修正,显式最小化属性误差,同时控制编辑复杂度和与原型的偏离。一个大规模自动构建的数据集包含片段编辑的推理链与属性变化量,支持两个阶段的确定性、可复现监督与可控多跳推理。相比先前工作,本框架通过片段建模更优地推理分子,并支持向目标值可控优化。在两类属性约束(QED、LogP、分子量与HOMO、LUMO)下的实验表明,生成分子在有效性和多属性目标精确满足方面持续领先于强基线模型与图算法。

原文摘要 · Abstract (English)

Generating molecules that satisfy precise numeric constraints over multiple physicochemical properties is critical and challenging. Although large language models (LLMs) are expressive, they struggle with precise multi-objective control and numeric reasoning without external structure and feedback. We introduce \textbf{M olGen}, a fragment-level, retrieval-augmented, two-stage framework for molecule generation under multi-property constraints. Stage I : Prototype generation: a multi-agent reasoner performs retrieval-anchored, fragment-level edits to produce a candidate near the feasible region. Stage II : RL-based fine-grained optimization: a fragment-level optimizer trained with Group Relative Policy Optimization (GRPO) applies one- or multi-hop refinements to explicitly minimize the property errors toward our target while regulating edit complexity and deviation from the prototype. A large, automatically curated dataset with reasoning chains of fragment edits and measured property deltas underpins both stages, enabling deterministic, reproducible supervision and controllable multi-hop reasoning. Unlike prior work, our framework better reasons about molecules by leveraging fragments and supports controllable refinement toward numeric targets. Experiments on generation under two sets of property constraints (QED, LogP, Molecular Weight and HOMO, LUMO) show consistent gains in validity and precise satisfaction of multi-property targets, outperforming strong LLMs and graph-based algorithms.

分子生成多属性约束强化学习智能体协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。