arXiv:2601.16038cs.AI2026-01中稿 · ML4Molecules 2025被引 2

用知识图谱增强大模型,让化学合成规划更准确可靠

Grounding Large Language Models in Reaction Knowledge Graphs for Synthesis Retrieval

  • 将化学反应路径检索转为自然语言转图查询任务
  • 使用对齐示例的一次提示效果最佳,零样本下自校验提升可执行性
  • 提供可复现的评估框架,适合做分子合成规划的研究者

大型语言模型(LLM)可用于辅助化学合成规划,但标准提示方法常产生幻觉或过时建议。本文将反应路径检索建模为文本到Cypher查询生成问题,定义了单步与多步检索任务。对比了零样本提示与基于静态、随机及嵌入选择的单样本变体,并引入检查清单驱动的验证与纠错循环。评估聚焦查询有效性与检索准确性。结果表明,使用对齐示例的一次提示表现最优;检查清单式自校正主要提升零样本设置下的可执行性,在已有优质示例时额外收益有限。我们提供了可复现的Text2Cypher评估设置,以推动基于知识图谱的LLM在合成规划中的研究。代码已开源:https://github.com/Intelligent-molecular-systems/KG-LLM-Synthesis-Retrieval。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can aid synthesis planning in chemistry, but standard prompting methods often yield hallucinated or outdated suggestions. We study LLM interactions with a reaction knowledge graph by casting reaction path retrieval as a Text2Cypher (natural language to graph query) generation problem, and define single- and multi-step retrieval tasks. We compare zero-shot prompting to one-shot variants using static, random, and embedding-based exemplar selection, and assess a checklist-driven validator/corrector loop. To evaluate our framework, we consider query validity and retrieval accuracy. We find that one-shot prompting with aligned exemplars consistently performs best. Our checklist-style self-correction loop mainly improves executability in zero-shot settings and offers limited additional retrieval gains once a good exemplar is present. We provide a reproducible Text2Cypher evaluation setup to facilitate further work on KG-grounded LLMs for synthesis planning. Code is available at https://github.com/Intelligent-molecular-systems/KG-LLM-Synthesis-Retrieval.

化学合成知识图谱大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。