arXiv:2601.15780cs.CV2026-01

用合成视频测试视觉语言模型的空间与情境感知能力

Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video

  • 设计合成视频数据集,评估模型对动作意图和空间关系的判断
  • 现有模型表现仅略高于随机水平,尤其在视角变换下角色识别差
  • 提供带稳定颜色提示的轻量级解决方案,适合研究视觉推理缺陷

视觉语言模型在依赖细微时空或几何线索时,空间推理能力仍不稳定。本文提出一个合成基准,检测两种互补能力:情境意识(判断互动是否具有危害性)与空间意识(追踪谁对谁做了什么,以及相对位置与运动)。通过极简视频对,测试三个挑战:区分暴力与非暴力行为、跨视角绑定攻击者角色、判断细粒度轨迹对齐。我们在无训练条件下评估近期VLMs,该基准适用于任何视频分类模型。结果表明,各任务表现仅略高于随机水平。引入简单辅助手段——稳定颜色线索,部分缓解攻击者角色混淆,但未能解决根本缺陷。我们发布数据与代码,旨在提供可复现的诊断工具,并推动轻量级空间先验的探索以补充大规模预训练。

原文摘要 · Abstract (English)

Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two complementary skills: situational awareness (recognizing whether an interaction is harmful or benign) and spatial awareness (tracking who does what to whom, and reasoning about relative positions and motion). Through minimal video pairs, we test three challenges: distinguishing violence from benign activity, binding assailant roles across viewpoints, and judging fine-grained trajectory alignment. While we evaluate recent VLMs in a training-free setting, the benchmark is applicable to any video classification model. Results show performance only slightly above chance across tasks. A simple aid, stable color cues, partly reduces assailant role confusions but does not resolve the underlying weakness. By releasing data and code, we aim to provide reproducible diagnostics and seed exploration of lightweight spatial priors to complement large-scale pretraining.

视觉语言模型空间推理合成数据情境感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。