构建复杂场景下时空定位新基准,解决模型在真实世界中的泛化短板。
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- 提出前后双向精修标注流程,提升标签质量
- 发现模型在复杂场景下平均性能下降10.4%
- 设计无需训练的两阶段框架,显著提升定位精度
时空视频定位(STVG)旨在根据自然语言描述定位视频中的目标对象。尽管多模态大模型取得进展,当前模型与真实世界需求之间仍存在显著差距,主要源于基准数据集范围有限,导致模型出现类别偏见、推理简化和语言鲁棒性差等问题。为此,我们推出OmniGround,一个涵盖3,475个视频、81个类别的综合性基准,支持复杂现实场景下的查询。我们提出前-后-精修标注流程,结合多向追踪与智能纠错机制,生成高质量标签。进一步引入DeepSTG评估框架,从四个互补维度量化数据集质量。评估显示,模型在复杂真实场景下平均性能下降10.4%,尤其在小目标、遮挡目标及复杂空间关系上表现更差。基于此,我们提出PG-TAF——一种无需训练的两阶段框架,将STVG分解为高层时间定位与细粒度时空传播。实验表明,PG-TAF在OmniGround上实现m_tIoU提升25.6%、m_vIoU提升35.6%,并在四个基准上保持一致增益。
原文摘要 · Abstract (English)
Spatio-Temporal Video Grounding (STVG) aims to localize target objects in videos based on natural language descriptions. Despite recent advances in Multimodal Large Language Models, a significant gap remains between current models and real-world demands involving diverse objects and complex queries. We attribute this to limited benchmark scope, causing models to exhibit category bias, oversimplified reasoning, and poor linguistic robustness. To address these limitations, we introduce OmniGround, a comprehensive benchmark with 3,475 videos spanning 81 categories and complex real-world queries. We propose the Forward-Backward-Refinement annotation pipeline that combines multi-directional tracking with intelligent error correction for high-quality labels. We further introduce DeepSTG, a systematic evaluation framework quantifying dataset quality across four complementary dimensions beyond superficial statistics. Evaluations reveal performance average drop of 10.4% on complex real-world scenes, particularly with small/occluded objects and intricate spatial relations. Motivated by these, we propose PG-TAF, a training-free two-stage framework decomposing STVG into high-level temporal grounding and fine-grained spatio-temporal propagation. Experiments demonstrate PG-TAF achieves 25.6% and 35.6% improvements in m\_tIoU and m\_vIoU on OmniGround with consistent gains across four benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。