构建首个大规模多实例时空动作定位基准,支持复杂场景下动作的精准定位与追踪。
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
- 提出时空动作定位新任务,要求模型同步识别、追踪并定位多主体动作。
- 构建含688视频、19,590标注的SVAG-Bench,平均每视频28.5个动作查询。
- 提供带人类验证的高质量标注流程与标准化评估工具,适合多智能体交互研究者。
真正具备能力的AI系统不仅应孤立地检测物体或识别行为,更需建立谁在何时何地执行何种动作的统一、具身化表征。这为高层推理、规划与具身交互提供感知基础。当前视频基准仅孤立评估部分能力,如空间定位、目标追踪或时间定位,无法衡量其联合集成进展。本文提出时空动作定位(SVAG)任务与配套基准SVAG-Bench,要求模型在复杂多主体场景中同时检测、追踪并时间定位满足自然语言查询的所有动作实例。该基准包含688段视频、19,590条经验证的标注及903个独特动作动词,源自城市、野生动物与交通监控场景,每视频平均含28.5个以动作为中心的查询,是同类基准中最密集的标注。标注通过专家人工标注、GPT-3.5改写增强与人工复核的流水线生成,确保语言多样性与准确性。我们还发布SVAGEval——标准化多指代评估工具包,并提出一种模块化基线架构SVAGFormer。
原文摘要 · Abstract (English)
A truly capable AI system must do more than detect objects or recognize activities in isolation. It must form unified, grounded representations of who is acting, what they are doing, and when and where these actions unfold. These representations provide the perceptual bedrock for high-level reasoning, planning, and embodied interaction in the real world. Building such agents is central to long-horizon goals in embodied AI and robotics. Current video benchmarks evaluate fragments of these capabilities in isolation. They focus on either spatial grounding, object tracking, or temporal localization. As a result, they cannot rigorously measure progress on their joint, multi-instance integration. We introduce Spatio-temporal Video Action Grounding (SVAG), a task and benchmark that explicitly targets this unified competence by requiring models to simultaneously detect, track, and temporally localize all objects that satisfy a natural language query in complex, multi-actor scenes. To support this task, we construct SVAG-Bench. It comprises 688 videos, 19,590 verified annotations, and 903 unique action verbs drawn from crowded urban environments, wildlife, and traffic surveillance. Each video has on average 28.5 action-centric queries. This yields the densest annotation among comparable video grounding benchmarks and enables fine-grained evaluation of multi-actor disambiguation, temporal overlap, and action compositionality. Annotations are produced by a pipeline that combines expert manual labeling, GPT-3.5 paraphrase augmentation, and human verification to ensure both linguistic diversity and correctness. We further release SVAGEval, a standardized multi-referent evaluation toolkit. We also introduce SVAGFormer, a strong modular baseline architecture for SVAG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。