arXiv:2508.11955cs.CV2025-08

用时间标注提升视频目标分割的语义对齐效果

Temporal Grounding as a Learning Signal for Referring Video Object Segmentation

  • 引入时间定位学习框架,显式建模语言与视频时序的对应关系
  • 在MeViS-M数据集上达到新最佳性能,提升显著
  • 适合关注视频理解与多模态对齐的研究者

指代式视频目标分割(RVOS)旨在根据自然语言描述分割并追踪视频中的物体,需精确对齐视觉内容与文本查询。现有方法常因随意采样帧及对所有可见物体进行无差别监督而产生语义错位。我们识别出根本问题在于传统训练范式缺乏明确的时间学习信号。为此,我们在具有挑战性的MeViS基准之上构建了MeViS-M数据集,手动标注每个物体被语言提及的时间跨度,提供直接、语义接地的监督信号。为利用该信号,我们提出时间定位学习(TGL)框架,包含两项关键策略:第一,时刻引导双路径传播(MDP),将相关时刻的语言引导分割与非相关时刻的语言无关传播解耦;第二,对象级选择性监督(OSS),仅在每段训练片段中监督与表达时间对齐的对象,减少语义噪声,强化语言条件学习。大量实验表明,我们的TGL框架有效利用时间信号,在挑战性的MeViS基准上建立新最优结果。代码与数据集将公开发布。

原文摘要 · Abstract (English)

Referring Video Object Segmentation (RVOS) aims to segment and track objects in videos based on natural language expressions, requiring precise alignment between visual content and textual queries. However, existing methods often suffer from semantic misalignment, largely due to indiscriminate frame sampling and supervision of all visible objects during training -- regardless of their actual relevance to the expression. We identify the core problem as the absence of an explicit temporal learning signal in conventional training paradigms. To address this, we introduce MeViS-M, a dataset built upon the challenging MeViS benchmark, where we manually annotate temporal spans when each object is referred to by the expression. These annotations provide a direct, semantically grounded supervision signal that was previously missing. To leverage this signal, we propose Temporally Grounded Learning (TGL), a novel learning framework that directly incorporates temporal grounding into the training process. Within this frame- work, we introduce two key strategies. First, Moment-guided Dual-path Propagation (MDP) improves both grounding and tracking by decoupling language-guided segmentation for relevant moments from language-agnostic propagation for others. Second, Object-level Selective Supervision (OSS) supervises only the objects temporally aligned with the expression in each training clip, thereby reducing semantic noise and reinforcing language-conditioned learning. Extensive experiments demonstrate that our TGL framework effectively leverages temporal signal to establish a new state-of-the-art on the challenging MeViS benchmark. We will make our code and the MeViS-M dataset publicly available.

视频分割多模态对齐时间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。