arXiv:2605.10230cs.LG2026-05

用片段级编辑提升分子优化,避免语言模型的幻觉问题

FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization

论文配图:FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization
图 1 · 摘自论文原文
  • 基于上下文感知的片段排序与替换,替代传统语言生成
  • 在多个数据集上超越大模型和图方法,0.6B小模型即达领先性能
  • 适合需要高效、可靠分子设计的药物研发人员

分子优化旨在通过微小结构修改改进分子性质,同时保持与原始分子的相似性。现有基于语言模型的方法通常将任务视为提示驱动的序列生成,但依赖自然语言会带来数据扩展瓶颈,易产生化学幻觉,并忽略片段效应的强上下文依赖性。我们提出FORGE,一种两阶段框架,将分子优化重新定义为上下文感知的局部编辑。第一阶段利用自动挖掘并验证的低至高阶编辑对,而非昂贵的人工文本标注,根据完整分子上下文中片段的属性贡献对候选片段进行排序,引入化学先验知识;第二阶段生成明确的片段替换。基于仅0.6B参数的紧凑语言模型,FORGE还能通过上下文示例适应未见过的黑箱目标函数。在Prompt-MolOpt、PMO-1k和ChemCoTBench三个基准上,FORGE持续优于先前方法,包括更大规模的语言模型和图神经网络方法。结果凸显了显式片段级监督作为更易获取、可扩展且无幻觉的替代方案的价值。

原文摘要 · Abstract (English)

Molecular optimization seeks to improve a molecule through small structural edits while preserving similarity to the starting compound. Recent language-model approaches typically treat this task as prompt-conditioned sequence generation. However, relying on natural language introduces an inherent data-scaling bottleneck, often leads to chemical hallucinations, and ignores the strong context dependence of fragment effects. We present FORGE, a two-stage framework that reformulates molecular optimization as context-aware local editing. By utilizing automatically mined, verified low-to-high edit pairs instead of expensive human text annotations, Stage 1 ranks candidate fragments by their property contribution under the full molecular context to inject chemical prior, and Stage 2 generates explicit fragment replacements. Built on a compact 0.6B language model, FORGE further adapts to unseen black-box objectives through in-context demonstrations. Across Prompt-MolOpt, PMO-1k and ChemCoTBench, FORGE consistently outperforms prior methods, including substantially larger language models and graph methods. These results highlight the value of explicit fragment-level supervision as a more easily obtainable, scalable, and hallucination-less alternative to natural language training.

分子生成片段编辑药物设计语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。