arXiv:2503.10500cs.CV2025-03被引 8

提出多目标时空视频定位新任务,可同时定位文本中所有提及对象及其互动关系。

OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding

  • 设计新任务OmniSTVG,支持任意数量目标及互动关系的时空定位。
  • 构建10,018段视频、10.2M帧的大规模基准BOSTVG,覆盖287类场景,每视频含1-10个目标。
  • 提出OmniTube模型,基于Transformer结构,适配多目标定位,性能表现优异。

本文提出一种新的时空多目标视频定位任务——OmniSTVG,旨在从视频中定位文本查询中提及的所有空间与时间目标。相比传统仅定位单一目标的STVG,OmniSTVG可定位任意数量的目标及其在查询中提及的交互对象,更具灵活性与实际应用价值。为推动该任务发展,我们构建了大规模基准BOSTVG,包含10,018段视频、共10.2M帧,覆盖287个类别,涵盖多样化场景。每段视频配以自由形式文本查询,目标数量为1至10个不等。所有视频均经人工精标注与反复校验,确保高质量。据我们所知,BOSTVG是目前首个且最大的OmniSTVG基准。为激励后续研究,我们提出OmniTube方法,受Transformer-based STVG启发,专为多目标定位设计,展现出良好效果。我们将发布基准、模型与结果,推动超越传统STVG的全面理解方向。

原文摘要 · Abstract (English)

In this paper, we propose spatio-temporal omni-object video grounding, dubbed OmniSTVG, a new STVG task that aims at localizing spatially and temporally all targets mentioned in the textual query from videos. Compared to classic STVG locating only a single target, OmniSTVG enables localization of not only an arbitrary number of text-referred targets but also their interacting counterparts in the query from the video, making it more flexible and practical in real scenarios for comprehensive understanding. In order to facilitate exploration of OmniSTVG, we introduce BOSTVG, a large-scale benchmark dedicated to OmniSTVG. Specifically, our BOSTVG consists of 10,018 videos with 10.2M frames and covers a wide selection of 287 classes from diverse scenarios. Each sequence in BOSTVG, paired with a free-form textual query, encompasses a varying number of targets ranging from 1 to 10. To ensure high quality, each video is manually annotated with meticulous inspection and refinement. To our best knowledge, BOSTVG is to date the first and the largest benchmark for OmniSTVG. To encourage future research, we introduce a simple yet effective approach, named OmniTube, which, drawing inspiration from Transformer-based STVG methods, is specially designed for OmniSTVG and demonstrates promising results. By releasing BOSTVG, we hope to go beyond classic STVG by locating every object appearing in the query for more comprehensive understanding, opening up a new direction for STVG. Our benchmark, model, and results will be released at https://github.com/JellyYao3000/OmniSTVG.

视频定位多目标时空理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。