用滑块谜题测试视觉语言模型的空间推理能力。
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
- 基于滑块谜题设计交互式多模态评测基准。
- 2D任务表现优于3D和文本,仍远低于人类水平。
- 适合评估模型空间规划与视觉对齐能力。
视觉语言模型(VLMs)在空间推理和视觉对齐方面表现不佳。为此,我们提出iVISPAR,一个用于评估VLM作为智能体进行空间推理能力的交互式多模态基准。iVISPAR基于滑块谜题的变体,该经典问题要求逻辑规划、空间意识和多步推理。基准支持视觉3D、2D及文本输入模态,可全面评估VLM的规划与推理能力。我们评估了多种前沿开源与闭源VLM,对比其表现,并提供最优路径解与人类基线以衡量任务复杂度与可行性。结果表明,尽管VLM在2D任务中表现优于3D或文本场景,但在复杂空间配置下仍显著落后于人类,凸显当前VLM在视觉对齐方面的持续挑战,揭示其难以达到人类级认知能力。项目网站:https://microcosm.ai/ivispar
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. \mbox{iVISPAR} is based on a variant of the sliding tile puzzle, a classic problem that demands logical planning, spatial awareness, and multi-step reasoning. The benchmark supports visual 3D, 2D, and text-based input modalities, enabling comprehensive assessments of VLMs' planning and reasoning skills. We evaluate a broad suite of state-of-the-art open-source and closed-source VLMs, comparing their performance while also providing optimal path solutions and a human baseline to assess the task's complexity and feasibility for humans. Results indicate that while VLMs perform better on 2D tasks compared to 3D or text-based settings, they struggle with complex spatial configurations and consistently fall short of human performance, illustrating the persistent challenge of visual alignment. This underscores critical gaps in current VLM capabilities, highlighting their limitations in achieving human-level cognition. Project website: https://microcosm.ai/ivispar
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。