arXiv:2603.25253cs.CLcs.AI2026-03被引 1

构建交互式化学结构解析基准,评估大模型的科学推理能力

MolQuest: A Benchmark for Agentic Evaluation of Abductive Reasoning in Chemical Structure Elucidation

  • 将分子结构解析设为多轮互动任务,要求模型规划实验并融合多种光谱数据
  • 顶尖模型准确率仅约50%,多数模型低于30%
  • 适合关注AI科学推理、智能实验设计的研究者

大语言模型(LLMs)在推动科学发现方面潜力巨大,但其在真实科研场景中动态推理能力的系统评估仍显不足。现有科学评测基准多采用静态单轮问答格式,难以衡量复杂科学任务中所需的多步迭代与实验交互能力。为此,我们提出MolQuest,一个基于真实化学实验数据的代理式分子结构解析评估框架。不同于现有数据集,MolQuest将分子结构解析建模为多轮交互任务,要求模型主动规划实验步骤、整合异构光谱信息(如NMR、MS),并迭代修正结构假设。该框架系统评估了LLMs在庞大复杂化学空间中的溯因推理与策略决策能力。实证结果表明,当前前沿模型在真实科学场景中存在显著局限:即使是最先进的模型准确率也仅约50%,多数模型表现低于30%。本工作提供可复现、可扩展的面向科学的LLM评估框架,揭示了当前模型在战略科学推理上的关键差距,为未来实现能主动参与科研过程的AI指明方向。

原文摘要 · Abstract (English)

Large language models (LLMs) hold considerable potential for advancing scientific discovery, yet systematic assessment of their dynamic reasoning in real-world research remains limited. Current scientific evaluation benchmarks predominantly rely on static, single-turn Question Answering (QA) formats, which are inadequate for measuring model performance in complex scientific tasks that require multi-step iteration and experimental interaction. To address this gap, we introduce MolQuest, a novel agent-based evaluation framework for molecular structure elucidation built upon authentic chemical experimental data. Unlike existing datasets, MolQuest formalizes molecular structure elucidation as a multi-turn interactive task, requiring models to proactively plan experimental steps, integrate heterogeneous spectral sources (e.g., NMR, MS), and iteratively refine structural hypotheses. This framework systematically evaluates LLMs' abductive reasoning and strategic decision-making abilities within a vast and complex chemical space. Empirical results reveal that contemporary frontier models exhibit significant limitations in authentic scientific scenarios: notably, even state-of-the-art (SOTA) models achieve an accuracy of only approximately 50%, while the performance of most other models remains below the 30% threshold. This work provides a reproducible and extensible framework for science-oriented LLM evaluation, our findings highlight the critical gap in current LLMs' strategic scientific reasoning, setting a clear direction for future research toward AI that can actively participate in the scientific process.

科学推理分子结构大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。