arXiv:2606.05847cs.AI2026-06

用智能体探索修复无效分子结构,保留原始设计意图。

Agentic Molecular Recovery via Molecule-Aware Exploration

论文配图:Agentic Molecular Recovery via Molecule-Aware Exploration
图 1 · 摘自论文原文
  • 通过分子感知的不匹配追踪与多候选路径探索,实现精准修复。
  • 在三个基线模型生成的20个无效分子上,恢复效果最优。
  • 适合需要高保真分子生成的药物研发人员使用。

基于大语言模型的文本引导分子生成常产生无效SMILES。我们提出,应从关注有效性的修复转向保持分子身份的恢复:不仅要恢复化学有效性,还需保留目标相关的结构特征,还原描述所暗示的分子身份。这一视角揭示了现有修正策略的局限性:事后修复会扭曲关键结构,纯大模型修正引入全局漂移,通用智能体修正受限于贪婪的单候选轨迹,即使使用可执行的RDKit编辑工具也难以突破。为此,我们提出AMREC,结合分子感知的不匹配追踪、扩展候选探索与轨迹级选择。在三个主干模型生成的20个无效ChEBI分子上,AMREC在结构、精确匹配和字符串级指标上均表现最佳。

原文摘要 · Abstract (English)

Text-guided molecular generation with LLMs often yields invalid SMILES. We argue that invalid drafts should be addressed through a shift from validity-oriented repair to identity-preserving molecular recovery: the objective is not only to restore chemical validity, but also to preserve target-relevant structural cues and recover the molecular identity implied by the description. This perspective reveals the limitations of existing correction strategies. Post-hoc repair can recover validity while distorting key structures, LLM-only correction can introduce unintended global drift, and generic agentic correction remains constrained by greedy single-candidate trajectories even when equipped with executable RDKit edit tools. To address these limitations, we propose AMREC, which couples molecule-aware mismatch tracking with expanded candidate exploration and trajectory-level selection. On invalid ChEBI-20 drafts from three backbone models, AMREC achieves the strongest overall recovery profile across structural, exact-match, and string-level metrics.

分子生成智能体化学信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。