新基准测试揭示视觉语言模型在空间推理上的显著短板。
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
- 构建1100张高复杂度真实图像数据集,评估模型空间感知与推理能力。
- 最强模型Gemini-2.5-Pro仅77.14%准确率,顺序生成任务仅30.00%。
- 适合研究视觉语言模型空间认知与人类智能差距的学者参考。
空间推理是人类认知的核心,使人能感知、理解并互动于物理世界。它依赖对空间结构和对象间关系的精细理解,是复杂推理与决策的基础。为探究当前视觉语言模型(VLMs)是否具备类似能力,我们提出Jigsaw-Puzzles,一个包含1,100张精心筛选的真实世界图像的新型基准,具有高空间复杂性。基于该数据集,设计五项任务,严格评估VLMs在空间感知、结构理解与推理方面的能力,并刻意减少对领域特定知识的依赖,以更纯粹地检验通用空间推理能力。我们在24个先进VLMs上进行综合评估。结果表明,即使最强模型Gemini-2.5-Pro也仅达77.14%整体准确率,且在顺序生成任务中仅30.00%准确率,远低于人类参与者超过90%的表现。这一持续存在的差距凸显了进一步发展的必要性,使Jigsaw-Puzzles成为推动VLMs空间推理研究的挑战性与诊断性基准。
原文摘要 · Abstract (English)
Spatial reasoning is a core component of human cognition, enabling individuals to perceive, comprehend, and interact with the physical world. It relies on a nuanced understanding of spatial structures and inter-object relationships, serving as the foundation for complex reasoning and decision-making. To investigate whether current vision-language models (VLMs) exhibit similar capability, we introduce Jigsaw-Puzzles, a novel benchmark consisting of 1,100 carefully curated real-world images with high spatial complexity. Based on this dataset, we design five tasks to rigorously evaluate VLMs' spatial perception, structural understanding, and reasoning capabilities, while deliberately minimizing reliance on domain-specific knowledge to better isolate and assess the general spatial reasoning capability. We conduct a comprehensive evaluation across 24 state-of-the-art VLMs. The results show that even the strongest model, Gemini-2.5-Pro, achieves only 77.14% overall accuracy and performs particularly poorly on the Order Generation task, with only 30.00% accuracy, far below the performance exceeding 90% achieved by human participants. This persistent gap underscores the need for continued progress, positioning Jigsaw-Puzzles as a challenging and diagnostic benchmark for advancing spatial reasoning research in VLMs. Our project page is at https://zesen01.github.io/jigsaw-puzzles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。