arXiv:2604.15994cs.AI2026-04中稿 · EMNLP

测试大模型对化学反应图的结构推理能力,发现其在复杂拓扑结构上表现严重下降。

ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams

  • 用化学反应图构建多层级推理任务,检验模型对拓扑结构的理解能力。
  • 24个模型在全局结构任务上性能比基础任务低超30%,暴露推理短板。
  • 适合关注视觉推理、科学图表理解的研究者和开发者参考。

多模态大语言模型在识别单一视觉元素和简单线性图谱方面表现优异,但在面对包含分支路径、汇聚流和环形依赖的复杂拓扑结构时,推理能力急剧下降,甚至无法准确计数终点。现有基准未能揭示这一差距,主要关注语义理解而非结构推理。我们提出ReactBench,一个基于真实化学反应图的基准,用于揭示模型在结构推理上的根本局限。这些科学图谱天然涵盖从线性链到环状图的多样结构,同时要求精确的局部识别与连贯的全局推理。该基准包含1,618对专家标注的问答对,覆盖四个层级的任务维度。对24个MLLM的广泛评估显示,锚点任务与整体结构推理任务之间性能差距超过30%。受控消融实验确认瓶颈在于推理而非感知。研究结果暴露了模型在结构理解上的根本缺陷,并为提升视觉推理指明方向。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning over simple linear diagrams. However, when faced with complex topological structures involving branching paths, converging flows, and cyclic dependencies, their reasoning capabilities degrade sharply, even on tasks as basic as counting endpoints. Existing benchmarks fail to probe this gap, focusing on semantic comprehension rather than structural reasoning. We introduce ReactBench, a benchmark that reveals fundamental limitations in structural reasoning through chemical reaction diagrams. These real-world scientific diagrams offer an ideal testbed because they naturally span diverse structures from linear chains to cyclic graphs, while requiring both precise local recognition and coherent global reasoning. Our benchmark comprises 1,618 expert-annotated QA pairs across four hierarchical task dimensions. Extensive evaluation across 24 MLLMs reveals a significant performance gap exceeding 30% between anchor-based tasks and holistic structural reasoning tasks. Controlled ablations confirm this bottleneck lies in reasoning, not perception. These findings expose a fundamental deficit in structural understanding and establish directions for advancing visual reasoning.

视觉推理化学图谱结构理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。