测试视觉语言模型的空间变形推理能力,发现多数模型表现不佳。
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
- 构建2D到3D的无限阶梯式空间变形评测框架
- 几乎所有模型在3D变形任务中均无法正确推理
- 适合研究模型空间认知能力的学者参考
人类天生具备在空间中构建和操作物体图像与结构的直观空间推理能力。近年来,研究者致力于赋予视觉-语言模型(VLMs)类似的空间推理能力。然而,这些模型是否真正理解并操控空间对象仍不明确。为此,我们提出一个全新的评估框架,用于评测VLM在空间变形推理任务中的表现。具体而言,我们从2D到3D构建了一个空间变形推理基准测试集,利用数据生成引擎可无限生成无数据泄露的评估问题对。我们从正向推理(给定操作序列,推导最终状态)和逆向推理(给定最终状态,还原操作序列)两个方向探索模型的能力。采用阶梯式竞赛模式,以变形步骤数作为等级划分标准,旨在探究模型推理能力的边界。有趣的是,评测结果表明,几乎没有任何模型展现出合理的空间变形推理能力。即使经过针对性训练和主流推理增强方法的改进,模型在3D空间变形推理任务上依然表现欠佳。
原文摘要 · Abstract (English)
Humans naturally possess the spatial reasoning ability to form and manipulate images and structures of objects in space. There is an increasing effort to endow Vision-Language Models (VLMs) with similar spatial reasoning capabilities. However, it remains unclear whether these models truly understand and manipulate spatial objects or not. To address this question, we propose a new evaluation framework aimed at assessing the performance of VLMs in spatial deformation reasoning tasks. Specifically, we construct a benchmark for spatial deformation reasoning from 2D to 3D. Leveraging our data engine, we can generate unlimited evaluation problem pairs with infinite steps, without any data leakage. We explore whether the model can effectively perform spatial deformation reasoning from two directions: forward reasoning (given the operations, find the final state) and reverse reasoning (given the final state, determine the operations). We adopt a ladder competition format, using the number of deformation steps as the level classification criterion, with the goal of exploring the boundaries of the model's deformation reasoning capabilities. Interestingly, the benchmarking results reveal that almost no model demonstrates plausible spatial deformation reasoning abilities. Furthermore, even after applying targeted training and mainstream reasoning enhancement methods, the models are still unable to perform well on 3D spatial deformation reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。