构建合成数据与课程学习框架,提升复杂视频定位能力
Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

- 用时空场景图控制难度,生成分级的合成查询数据
- 现有模型在复合查询上性能骤降,平均分掩盖真实短板
- 提出课程强化学习框架,对难题提升显著
尽管近期多模态大模型在时空视频定位(STVG)任务上取得显著进展,但现有评估和训练数据主要关注简单查询,忽视了真实场景中常见的复合查询——目标需通过联合推理其属性与与其他实体的关系来消歧。为此,我们提出复合时空视频定位(CompSTVG)任务,要求模型处理包含交织属性与关系线索的复杂文本查询。为规模化推进该任务,我们构建了一个合成数据引擎,利用时空场景图作为难度度量,将可控难度的查询生成建模为约束规划问题,生成用于评估与训练的分级数据。基于此引擎,我们推出STVG-CompBench基准,按显式难度层级划分,同时涵盖时间复杂度与空间干扰。在该基准上评估11个代表性STVG模型发现,当前模型在复合查询上表现不佳,性能下降明显,但常被整体数据集平均分掩盖。我们进一步构建合成训练数据,并提出CurrSTVG课程强化学习框架,带来持续提升,尤其在最困难的复合查询上效果最显著。
原文摘要 · Abstract (English)
Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely overlook the compositional queries prevalent in real-world scenarios, where a target must be disambiguated by jointly reasoning about its attributes and relations to other entities. To bridge this gap, we propose Compositional Spatio-Temporal Video Grounding (CompSTVG), a task that requires models to process complex textual queries where every intertwined attribute and relational cue is essential for disambiguation. To facilitate this task at scale, we build a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training. Built on this engine, we introduce STVG-CompBench, a benchmark stratified by explicit difficulty levels that jointly capture temporal complexity and spatial interference. Evaluating 11 representative STVG models on STVG-CompBench reveals that current models perform poorly on compositional queries, exhibiting a sharp performance drop that is typically obscured by overall dataset-level averages. We further construct synthetic training data and propose CurrSTVG, a curriculum reinforcement learning framework that delivers consistent gains, with the largest improvements observed on the most challenging compositional queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。