arXiv:2503.13983cs.CV2025-03AAAI被引 32

让多模态大模型同时精准定位视频的时间和空间位置。

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

  • 用交错式时空感知查询捕捉动态时空信息
  • 在11个基准上达到领先性能,最高提升12.3%
  • 适合需要视频精确定位的研究与应用

多模态大语言模型在时间或空间定位上已取得显著进展,但在时空视频定位方面仍表现不佳。主要源于两大挑战:难以准确提取每帧的时空信息,以及视觉标记数量庞大导致难以精确映射到空间坐标。为此,我们提出SpaceVLLM,一种具备时空视频定位能力的多模态大模型。通过一组交错的时空感知查询捕捉时间感知与动态空间信息,并设计查询引导的空间解码器建立查询与空间坐标的对应关系。此外,由于缺乏时空数据集,我们构建了包含480K样本的统一时空定位(Uni-STG)数据集,覆盖三个任务。大量实验证明,SpaceVLLM在11个涵盖时间、空间、时空及视频理解任务的基准上均达到当前最优表现,充分展现了方法的有效性。代码、数据集与模型将开源。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have made remarkable progress in either temporal or spatial localization. However, they struggle to perform spatio-temporal video grounding. This limitation stems from two major challenges. Firstly, it is difficult to extract accurate spatio-temporal information of each frame in the video. Secondly, the substantial number of visual tokens makes it challenging to precisely map visual tokens of each frame to their corresponding spatial coordinates. To address these issues, we introduce SpaceVLLM, a MLLM endowed with spatio-temporal video grounding capability. Specifically, we adopt a set of interleaved Spatio-Temporal Aware Queries to capture temporal perception and dynamic spatial information. Moreover, we propose a Query-Guided Space Decoder to establish a corresponding connection between the queries and spatial coordinates. Additionally, due to the lack of spatio-temporal datasets, we construct the Unified Spatio-Temporal Grounding (Uni-STG) dataset, comprising 480K instances across three tasks. This dataset fully exploits the potential of MLLM to simultaneously facilitate localization in both temporal and spatial dimensions. Extensive experiments demonstrate that SpaceVLLM achieves the state-of-the-art performance across 11 benchmarks covering temporal, spatial, spatio-temporal and video understanding tasks, highlighting the effectiveness of our approach. Our code, datasets and model will be released at https://github.com/Jayce1kk/SpaceVLLM.

视频定位多模态时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。