arXiv:2508.17180cs.AIcs.CV2025-08被引 5

测试视觉图像中的数学与空间推理能力,发现当前模型表现不佳。

MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes

  • 构建新基准,评估模型从图像中进行数学推理的能力。
  • 模型在特征计数和变换识别任务上表现差,多依赖表面规律。
  • 适合研究多模态模型深层推理的学者使用。

多模态大语言模型(MLLMs)的一个关键前沿是直接从图像中执行深度数学与空间推理,超越其在语义描述上的成功。数学曲面图提供了一个严格的测试平台,可剥离自然图像中的语义噪声。为此,我们提出 MaRVL-QA(Mathematical Reasoning over Visual Landscapes),一个用于定量评估核心推理能力的新基准。该基准包含两个新任务:拓扑计数(识别并统计局部极大值等特征)和变换识别(判断应用的几何变换)。数据由经过严格歧义过滤的函数库生成。在 MaRVL-QA 上的评估显示,即使最先进的 MLLMs 也表现显著不足,常依赖表面启发式而非稳健的空间推理。MaRVL-QA 为研究社区提供了挑战性工具,可用于衡量进展、揭示模型局限,并引导具备更强推理能力的 MLLMs 发展。

原文摘要 · Abstract (English)

A key frontier for Multimodal Large Language Models (MLLMs) is the ability to perform deep mathematical and spatial reasoning directly from images, moving beyond their established success in semantic description. Mathematical surface plots provide a rigorous testbed for this capability, as they isolate the task of reasoning from the semantic noise common in natural images. To measure progress on this frontier, we introduce MaRVL-QA (Mathematical Reasoning over Visual Landscapes), a new benchmark designed to quantitatively evaluate these core reasoning skills. The benchmark comprises two novel tasks: Topological Counting, identifying and enumerating features like local maxima; and Transformation Recognition, recognizing applied geometric transformations. Generated from a curated library of functions with rigorous ambiguity filtering, our evaluation on MaRVL-QA reveals that even state-of-the-art MLLMs struggle significantly, often resorting to superficial heuristics instead of robust spatial reasoning. MaRVL-QA provides a challenging new tool for the research community to measure progress, expose model limitations, and guide the development of MLLMs with more profound reasoning abilities.

多模态数学推理视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。