首个评估视觉语言模型时空推理能力的基准,揭示现有模型在动态理解上的严重短板。
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- 构建包含真实与合成视频的多场景问答数据集,聚焦运动与视角变化
- 实测主流VLM在动态任务上远低于人类表现,尤其在多线索融合与时间连贯性上薄弱
- 提出4D特征场重建与时空微调等方向,为提升动态感知提供新思路
视觉语言模型(VLMs)在融合语言与视觉推理方面展现出显著能力,但对动态时空交互的理解仍存在根本性局限。人类能轻松追踪物体的移动、旋转和视角变化——这些能力对真实世界中的动态理解至关重要,却在当前的VLM中明显缺失。本文提出VLM4D,首个专门用于评估VLM时空推理能力的基准。该基准包含多样化的现实世界与合成视频,搭配精心设计的问题-答案对,强调平移与旋转运动、视角意识及运动连续性。通过对前沿开源与闭源VLM的全面评测,我们发现其性能显著落后于人类基线,暴露出现有模型在整合多重视觉线索与维持时间连贯性方面的根本缺陷。深入分析表明,引入4D特征场重建和针对性的时空监督微调可有效提升模型的时空理解能力。本研究旨在推动对VLM空间与时间定位能力的进一步探索,为动态环境下的更强大、可靠的视觉智能铺平道路。
原文摘要 · Abstract (English)
Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。