arXiv:2510.16590cs.LGcs.AI2025-10被引 3

用原子标识锚定思维链,让大模型高效完成化学逆合成任务。

Atom-anchored LLMs speak Chemistry: A Retrosynthesis Demonstration

  • 通过原子唯一标识将大模型推理过程与分子结构绑定,实现零样本和少样本推理。
  • 在药物分子上达到90%反应位点、74%最终原料的准确率,显著超越以往表现。
  • 适合缺乏标注数据的化学研究场景,尤其适合逆合成与分子设计领域专家。

机器学习在化学中的应用常受限于标注数据稀缺和成本高昂,传统监督方法难以推广。本文提出一种基于通用大语言模型(LLMs)的分子推理框架,无需特定任务训练。方法通过独特原子标识将思维链推理锚定在分子结构上:首先执行零样本任务,识别相关片段及其化学标签或转化类别;可选第二步中,利用少量示例进行少样本任务,预测具体化学转化。该框架应用于单步逆合成任务,此前大模型表现不佳。在学术基准和专家验证的药物发现分子上,成功率达90%以上识别化学上合理的反应位点,40%以上正确命名反应类型,74%以上准确预测最终反应物。本工作为分子推理与转化类问题提供了通用解决方案,确立了原子锚定式大模型在数据稀缺化学领域的强大潜力。

原文摘要 · Abstract (English)

Applications of machine learning in chemistry are often limited by the scarcity and expense of labeled data, restricting traditional supervised methods. In this work, we introduce a framework for molecular reasoning using general-purpose Large Language Models (LLMs) that operates without requiring task-specific model training. Our method anchors chain-of-thought reasoning to the molecular structure by using unique atomic identifiers. First, the LLM performs a zero-shot task to identify relevant fragments and their associated chemical labels or transformation classes. In an optional second step, this position-aware information is used in a few-shot task with provided class examples to predict the chemical transformation. We apply our framework to single-step retrosynthesis, a task where LLMs have previously underperformed. Across academic benchmarks and expert-validated drug discovery molecules, our work enables LLMs to achieve high success rates in identifying chemically plausible reaction sites ($\geq90\%$), named reaction classes ($\geq40\%$), and final reactants ($\geq74\%$). Ultimately, our work establishes a general blueprint for applying LLMs to challenges where molecular reasoning and molecular transformations are key, positioning atom-anchored LLMs as a powerful solution for data-scarce chemistry domains.

逆合成大模型分子推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。