arXiv:2505.07735cs.LG2025-05被引 26

最新大模型可直接完成有机化学推理,准确率达50%以上。

Assessing the Chemical Intelligence of Large Language Models

  • 构建新基准ChemIQ,要求模型生成短答案而非选择题。
  • 顶级模型在高推理模式下正确率50%-57%,远超非推理模型(3%-7%)。
  • 能解析NMR数据并生成SMILES,部分任务接近人类水平。

大型语言模型是多功能通用工具,近年来的“推理模型”在数学与软件工程等高级问题解决领域表现显著提升。本文评估了推理模型在无外部工具辅助下直接完成化学任务的能力。我们创建了新基准ChemIQ,包含816道题目,聚焦有机化学中的分子理解与化学推理,采用短答案形式,更贴近真实应用。在最高推理模式下,OpenAI o3-mini、Google Gemini Pro 2.5和DeepSeek R1正确率分别为50%-57%,且更高推理层级显著提升性能。相比而言,非推理模型仅达3%-7%准确率。模型首次实现将SMILES字符串转换为IUPAC名称。此外,最新推理模型可从一维和二维1H、13C NMR数据中推断结构,Gemini Pro 2.5对含最多10个重原子的分子正确生成SMILES比例达约90%,甚至成功解析含25个重原子的复杂结构。各任务中,模型推理过程与人类化学家高度相似。

原文摘要 · Abstract (English)

Large Language Models are versatile, general-purpose tools with a wide range of applications. Recently, the advent of "reasoning models" has led to substantial improvements in their abilities in advanced problem-solving domains such as mathematics and software engineering. In this work, we assessed the ability of reasoning models to perform chemistry tasks directly, without any assistance from external tools. We created a novel benchmark, called ChemIQ, consisting of 816 questions assessing core concepts in organic chemistry, focused on molecular comprehension and chemical reasoning. Unlike previous benchmarks, which primarily use multiple choice formats, our approach requires models to construct short-answer responses, more closely reflecting real-world applications. The reasoning models, OpenAI's o3-mini, Google's Gemini Pro 2.5, and DeepSeek R1, answered 50%-57% of questions correctly in the highest reasoning modes, with higher reasoning levels significantly increasing performance on all tasks. These models substantially outperformed the non-reasoning models which achieved only 3%-7% accuracy. We found that Large Language Models can now convert SMILES strings to IUPAC names, a task earlier models were unable to perform. Additionally, we show that the latest reasoning models can elucidate structures from 1D and 2D 1H and 13C NMR data, with Gemini Pro 2.5 correctly generating SMILES strings for around 90% of molecules containing up to 10 heavy atoms, and in one case solving a structure comprising 25 heavy atoms. For each task, we found evidence that the reasoning process mirrors that of a human chemist. Our results demonstrate that the latest reasoning models can, in some cases, perform advanced chemical reasoning.

化学智能推理模型NMR解析SMILES

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。