arXiv:2507.18140cs.CL2025-07被引 4

评测大模型在数学推理中基于代码的图像操作能力,填补视觉细粒度操作评估空白。

MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning

  • 用代码作为中间表示,评估模型从零生成和编辑数学图像的能力。
  • 在五类常见数学图上测试,模型在细粒度操作上仍远低于人类水平。
  • 适合关注多模态模型视觉推理与数学能力的研究者使用。

多模态大语言模型(MLLMs)在数学推理中已能基于文本指令执行视觉操作。现有评估主要关注纯文本输出,忽视了模型通过代码实现精准视觉操作的能力。本文提出首个针对该能力的细粒度评估基准MathOPEval,聚焦两大核心任务:(1) 多模态代码生成(MCG),评估模型从零构建可视化的能力;(2) 多模态代码编辑(MCE),评估删除、修改、标注三类细粒度操作。评估覆盖几何图、函数图像及三类统计图表,共五类主流数学图。实验涵盖九个主流MLLMs,结果表明当前模型在细粒度视觉操作方面仍显著落后于人类表现。

原文摘要 · Abstract (English)

Recent progress in Multi-modal Large Language Models (MLLMs) has enabled step-by-step multi-modal mathematical reasoning by performing visual operations based on the textual instructions. A promising approach uses code as an intermediate representation to precisely express and manipulate the images in the reasoning steps. However, existing evaluations focus mainly on text-only reasoning outputs, leaving the MLLM's ability to perform accurate visual operations via code largely unexplored. This work takes a first step toward addressing that gap by evaluating MLLM's code-based capabilities in multi-modal mathematical reasoning.Specifically, our framework focuses on two key evaluation aspects: (1) Multi-modal Code Generation (MCG) evaluates the model's ability to accurately understand and construct visualizations from scratch. (2) Multi-modal Code Editing (MCE) assesses the model's capacity for fine-grained operations, which include three types: Deletion, Modification and Annotation. To evaluate the above tasks, we incorporate a dataset that covers the five most popular types of mathematical figures, including geometric diagrams, function plots, and three types of statistical charts, to provide a comprehensive and effective measurement of existing MLLMs. Our experimental evaluation involves nine mainstream MLLMs, and the results reveal that existing models still lag significantly behind human performance in performing fine-grained visual operations.

多模态数学推理代码生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。