arXiv:2510.22340cs.AIcs.CL2025-10被引 1

首个动态立体几何推理基准,评估视觉语言模型真实空间思维能力

DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry

  • 通过半自动流程构建可动态生成的立体几何题库
  • 发现主流模型在动态场景下性能大幅下降,空间推理能力薄弱
  • 兼顾答案与推理过程评估,适合研究空间智能的学者

立体几何求解需要融合空间智能与符号推理的空间数学思维。然而,现有多模态数学推理基准主要关注二维平面几何,依赖静态数据集易发生数据污染和记忆现象,且仅以最终答案评价模型表现,忽视推理过程。为此,我们提出DynaSolidGeo,首个用于评估视觉语言模型(VLMs)真实空间推理能力的动态基准。该基准通过半自动标注流程构建,包含503个专家精选的初始问题,理论上可动态生成无限多样本的图文实例。除答案准确率外,还引入基于专家标注推理链的过程评估,衡量逻辑有效性与因果连贯性。在代表性开源与闭源VLM上的实验显示:存在显著性能差距,动态环境下严重退化,尤其在需高阶空间智能的任务(如心理旋转、可视化)上表现不佳。代码与数据集已公开于DynaSolidGeo。

原文摘要 · Abstract (English)

Solid geometry problem solving demands spatial mathematical reasoning that integrates spatial intelligence and symbolic reasoning. However, most existing multimodal mathematical reasoning benchmarks focus primarily on 2D plane geometry, rely on static datasets prone to data contamination and memorization, and evaluate models solely by final answers, overlooking the reasoning process. To address these limitations, we introduce DynaSolidGeo, the first dynamic benchmark for evaluating genuine spatial reasoning in Vision-Language Models (VLMs). Constructed through a semi-automatic annotation pipeline, DynaSolidGeo contains 503 expert-curated seed questions that can, in principle, dynamically generate an unbounded number of diverse multimodal text-visual instances. Beyond answer accuracy, we incorporate process evaluation based on expert-annotated reasoning chains to measure logical validity and causal coherence. Experiments across representative open-source and closed-source VLMs reveal large performance gaps, severe degradation in dynamic settings, and poor performance on tasks requiring high-level spatial intelligence, such as mental rotation and visualization. The code and dataset are available at \href{https://zgca-ai4edu.github.io/DynaSolidGeo/}{DynaSolidGeo}.

空间推理视觉语言模型动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。