arXiv:2601.01547cs.CVcs.AI2026-01被引 1

提出新评估框架,发现视觉语言模型在物理推理上远落后于人类。

Vision-language models lag human performance on physical dynamics and intent reasoning

  • 构建了包含1.1万段真实视频的EscherVerse数据集,用于测试目标导向的空间推理能力。
  • 最先进模型准确率仅57.26%,人类平均达90.62%,差距显著。
  • 适合关注具身智能、空间推理与人类认知对比的研究者。

空间智能是具身认知的核心,但当前AI系统仍难以理解开放世界中的人类环境里的物理交互。尽管在受控基准上表现良好,视觉语言模型往往无法联合建模物理动态、参考系及驱动空间变化的潜在人类意图。我们提出了目标-空间智能(Teleo-Spatial Intelligence, TSI),将时空变化与目标导向结构关联。为评估TSI,我们构建了EscherVerse,一个基于11,328段真实视频的大规模开放世界资源,包含8,000个评测样本和35,963个指令微调样本。在27个主流视觉语言模型和11名标注者的第一轮人类反应独立分析中,我们发现持续存在的目标-空间推理差距:最强专有模型总体准确率为57.26%,远低于人类第一轮表现(84.81%至95.14%,均值90.62%)。在真实世界、意图感知数据上微调可缩小开放权重模型的差距,但无法消除。EscherVerse提供了一个目的感知的空间推理诊断平台,凸显了模式识别与人类级理解之间的关键鸿沟。

原文摘要 · Abstract (English)

Spatial intelligence is central to embodied cognition, yet contemporary AI systems still struggle to reason about physical interactions in open-world human environments. Despite strong performance on controlled benchmarks, vision-language models often fail to jointly model physical dynamics, reference frames, and the latent human intentions that drive spatial change. We introduce Teleo-Spatial Intelligence (TSI), a reasoning capability that links spatiotemporal change to goal-directed structure. To evaluate TSI, we present EscherVerse, a large-scale open-world resource built from 11,328 real-world videos, including an 8,000-example benchmark and a 35,963-example instruction-tuning set. Across 27 state-of-the-art vision-language models and an independent analysis of first-pass human responses from 11 annotators, we identify a persistent teleo-spatial reasoning gap: the strongest proprietary model achieves 57.26\% overall accuracy, far below first-pass human performance, which ranges from 84.81\% to 95.14\% with a mean of 90.62\%. Fine-tuning on real-world, intent-aware data narrows this gap for open-weight models, but does not close it. EscherVerse provides a diagnostic testbed for purpose-aware spatial reasoning and highlights a critical gap between pattern recognition and human-level understanding in embodied AI.

具身智能空间推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。