arXiv:2601.15279cs.LGcs.AI2026-01被引 5

用符号验证法评估模型对分子结构的推理能力

MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs

  • 基于分子图的符号可验证任务,精准测试模型结构理解力
  • 发现现有模型在特定分子结构上存在系统性失败
  • 为化学大模型研发提供可操作的改进方向

分子的性质由其组成与结构决定,这些信息编码于分子图中。因此,对分子性质的推理需要解析和理解分子图的能力。大型语言模型(LLMs)在化学领域应用日益广泛,涵盖分子命名转换、描述生成、文本引导生成及性质或反应预测等任务。然而,现有基准大多依赖文献或代理标签,存在泄露或偏差风险,或仅以选择题形式评估。我们提出MolecularIQ,一个专注于符号可验证任务的分子结构推理基准。该基准支持对分子图推理能力的细粒度评估,揭示了模型在特定任务和分子结构上的能力模式,帮助定位模型缺陷,为当前化学大模型的局限性提供可行动的洞察,并指导开发能忠实推理分子结构的模型。

原文摘要 · Abstract (English)

A molecule's properties are fundamentally determined by its composition and structure encoded in its molecular graph. Thus, reasoning about molecular properties requires the ability to parse and understand the molecular graph. Large Language Models (LLMs) are increasingly applied to chemistry, tackling tasks such as molecular name conversion, captioning, text-guided generation, and property or reaction prediction. Most existing benchmarks emphasize general chemical knowledge, rely on literature or surrogate labels that risk leakage or bias, or reduce evaluation to multiple-choice questions. We introduce MolecularIQ, a molecular structure reasoning benchmark focused exclusively on symbolically verifiable tasks. MolecularIQ enables fine-grained evaluation of reasoning over molecular graphs and reveals capability patterns that localize model failures to specific tasks and molecular structures. This provides actionable insights into the strengths and limitations of current chemistry LLMs and guides the development of models that reason faithfully over molecular structure.

分子推理符号验证化学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。