测试大模型在多步视觉模拟任务中的空间认知能力
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
- 构建新基准STARE,评估模型对几何变换与空间推理的视觉模拟能力
- 3D立方体展开和七巧板等复杂任务上模型表现接近随机
- 适合关注多模态模型视觉推理缺陷的研究者和开发者
空间认知是人类智能的核心,使人能通过视觉模拟解决问题而非仅依赖语言推理。现有AI评估基准主要考察语言推理,忽视了非语言、多步骤视觉模拟的复杂性。我们提出STARE(空间变换与推理评估)基准,用于严格评估多模态大语言模型在更需多步视觉模拟的任务上的表现。该基准包含4000个任务,涵盖基础几何变换(2D与3D)、整合空间推理(立方体展开、七巧板拼图)及真实世界空间推理(视角与时间推理),反映物体组装、机械图解读和日常空间导航等实际认知挑战。评估显示,模型在简单2D变换上表现良好,但在复杂任务如3D立方体展开和七巧板拼图上表现接近随机(准确率未达显著水平)。人类在复杂任务中接近完美准确率,但耗时较长(最长28.9秒),借助中间视觉模拟可平均提速7.5秒。相比之下,模型在视觉模拟辅助下表现不一:多数任务有提升,但在特定任务如七巧板(GPT-4o, o1)和立方体展开(Claude-3.5, Gemini-2.0 Flash)中反而下降,表明模型可能无法有效利用中间视觉信息。
原文摘要 · Abstract (English)
Spatial cognition is essential for human intelligence, enabling problem-solving through visual simulations rather than solely relying on verbal reasoning. However, existing AI benchmarks primarily assess verbal reasoning, neglecting the complexities of non-verbal, multi-step visual simulation. We introduce STARE(Spatial Transformations and Reasoning Evaluation), a benchmark designed to rigorously evaluate multimodal large language models on tasks better solved through multi-step visual simulation. STARE features 4K tasks spanning foundational geometric transformations (2D and 3D), integrated spatial reasoning (cube net folding and tangram puzzles), and real-world spatial reasoning (perspective and temporal reasoning), reflecting practical cognitive challenges like object assembly, mechanical diagram interpretation, and everyday spatial navigation. Our evaluations show that models excel at reasoning over simpler 2D transformations, but perform close to random chance on more complex tasks like 3D cube net folding and tangram puzzles that require multi-step visual simulations. Humans achieve near-perfect accuracy but take considerable time (up to 28.9s) on complex tasks, significantly speeding up (down by 7.5 seconds on average) with intermediate visual simulations. In contrast, models exhibit inconsistent performance gains from visual simulations, improving on most tasks but declining in specific cases like tangram puzzles (GPT-4o, o1) and cube net folding (Claude-3.5, Gemini-2.0 Flash), indicating that models may not know how to effectively leverage intermediate visual information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。