首个面向有机机理推理的权威评测基准,助力AI真正理解化学反应路径。
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
- 构建超万条专家标注的机理步骤数据集,支持精细评估
- 发现主流大模型在多步推理中一致性差,准确率不足
- 提出动态评分框架oMeS,适配化学逻辑与相似性双重验证
有机反应机理是反应物经由一系列基元步骤生成中间体和产物的过程,是理解化学反应活性及设计新分子与反应的基础。尽管大语言模型(LLMs)在合成设计等化学任务中展现出潜力,但其是否具备真正的化学推理能力——即生成合理中间体、保持化学一致性、遵循逻辑连贯的多步路径——仍不明确。为此,我们提出了oMeBench,首个大规模、专家标注的有机机理推理评测基准,包含超过10,000条带有中间体、类型标签和难度评级的机理步骤。为进一步精确评估并实现细粒度打分,我们提出oMeS动态评估框架,结合步骤级逻辑与化学相似性。对当前先进LLMs的分析显示,尽管模型展现出一定的化学直觉,但在正确且一致的多步推理方面仍表现不佳。值得注意的是,采用提示策略并基于本数据集微调专用模型,性能比领先闭源模型提升50%。我们期望oMeBench能为推动AI向真实化学推理迈进提供严谨基础。
原文摘要 · Abstract (English)
Organic reaction mechanisms are the stepwise elementary reactions by which reactants form intermediates and products, and are fundamental to understanding chemical reactivity and designing new molecules and reactions. Although large language models (LLMs) have shown promise in understanding chemical tasks such as synthesis design, it is unclear to what extent this reflects genuine chemical reasoning capabilities, i.e., the ability to generate valid intermediates, maintain chemical consistency, and follow logically coherent multi-step pathways. We address this by introducing oMeBench, the first large-scale, expert-curated benchmark for organic mechanism reasoning in organic chemistry. It comprises over 10,000 annotated mechanistic steps with intermediates, type labels, and difficulty ratings. Furthermore, to evaluate LLM capability more precisely and enable fine-grained scoring, we propose oMeS, a dynamic evaluation framework that combines step-level logic and chemical similarity. We analyze the performance of state-of-the-art LLMs, and our results show that although current models display promising chemical intuition, they struggle with correct and consistent multi-step reasoning. Notably, we find that using prompting strategy and fine-tuning a specialist model on our proposed dataset increases performance by 50% over the leading closed-source model. We hope that oMeBench will serve as a rigorous foundation for advancing AI systems toward genuine chemical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。