arXiv:2505.12702cs.CV2025-05被引 6

构建首个面向长视频的指代分割基准,推动真实场景下视频理解发展。

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

  • 提出长时序指代视频对象分割新基准Long-RVOS,覆盖60秒以上长视频。
  • 引入时序与时空一致性评估指标,突破传统逐帧评价局限。
  • 设计简洁有效的ReferMo模型,显著提升长视频中运动与依赖建模能力。

指代视频对象分割(RVOS)旨在根据语言描述识别、追踪并分割视频中的目标物体,近年来受到广泛关注。然而现有数据集仍局限于数秒内的短片段,且多数帧中目标物体明显可见。为推动该任务向更实际场景演进,本文提出 extbf{Long-RVOS},一个大规模长时序指代视频对象分割基准。Long-RVOS 包含超过2,000段视频,平均时长超过60秒,涵盖多种物体经历遮挡、消失-重现和镜头切换的情况。物体由人工标注三类不同描述,分别用于评估对静态属性、运动模式及时空关系的理解。此外,不同于以往仅依赖逐帧空间评估的基准,本文引入两个新指标以评估时序与时空一致性。在 Long-RVOS 上对6个先进方法进行基准测试,结果表明当前方法在长视频挑战下表现严重受限。为此,我们进一步提出 ReferMo,一种融合运动信息以扩展时序感受野,并采用局部到全局架构捕捉短期动态与长期依赖的基线方法。尽管结构简单,ReferMo 在长时序场景下相比现有方法实现显著提升。我们希望 Long-RVOS 与基线能推动未来 RVOS 研究迈向更真实、更长视频的任务。

原文摘要 · Abstract (English)

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible in most frames. To advance the task towards more practical scenarios, we introduce \textbf{Long-RVOS}, a large-scale benchmark for long-term referring video object segmentation. Long-RVOS contains 2,000+ videos of an average duration exceeding 60 seconds, covering a variety of objects that undergo occlusion, disappearance-reappearance and shot changing. The objects are manually annotated with three different types of descriptions to individually evaluate the understanding of static attributes, motion patterns and spatiotemporal relationships. Moreover, unlike previous benchmarks that rely solely on the per-frame spatial evaluation, we introduce two new metrics to assess the temporal and spatiotemporal consistency. We benchmark 6 state-of-the-art methods on Long-RVOS. The results show that current approaches struggle severely with the long-video challenges. To address this, we further propose ReferMo, a promising baseline method that integrates motion information to expand the temporal receptive field, and employs a local-to-global architecture to capture both short-term dynamics and long-term dependencies. Despite simplicity, ReferMo achieves significant improvements over current methods in long-term scenarios. We hope that Long-RVOS and our baseline can drive future RVOS research towards tackling more realistic and long-form videos.

视频分割长视频语言理解基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。