arXiv:2510.18873cs.CV2025-10被引 19

构建动态空间智能评测基准,揭示视觉模型在运动推理中的关键缺陷

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence

  • 设计9类解耦运动模式,通过近1000段动态视频评估模型
  • 14个模型普遍混淆观察者与物体运动,相对关系推断准确率低
  • 适合研究动态场景理解、具身智能的学者和开发者参考

动态空间关系推理至关重要,因观察者与物体常同时移动。尽管视觉语言模型(VLMs)和视觉专家模型在二维静态任务中表现优异,但其对动态三维场景的理解能力仍有限。我们提出动态空间智能,并构建DSI-Bench,包含近1,000段动态视频及超过1,700个手动标注问题,覆盖九种解耦的观察者与物体运动模式。空间-时间对称设计减少偏差,实现对自运动与物体运动推理的系统评估。对14个VLMs和专家模型的评估显示:模型常混淆观察者与物体运动,存在语义偏差,且在动态场景中难以准确推断相对关系。DSI-Bench为通用模型与专业模型未来发展中动态空间智能的研究提供了关键洞见。

原文摘要 · Abstract (English)

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise models excel in 2D tasks and static scenarios, their ability to fully understand dynamic 3D scenarios remains limited. We introduce Dynamic Spatial Intelligence and propose DSI-Bench, a benchmark with nearly 1,000 dynamic videos and over 1,700 manually annotated questions covering nine decoupled motion patterns of observers and objects. Spatially and temporally symmetric designs reduce biases and enable systematic evaluation of models' reasoning about self-motion and object motion. Our evaluation of 14 VLMs and expert models reveals key limitations: models often conflate observer and object motion, exhibit semantic biases, and fail to accurately infer relative relationships in dynamic scenarios. Our DSI-Bench provides valuable findings and insights about the future development of general and expertise models with dynamic spatial intelligence.

动态推理视觉模型空间智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。