arXiv:2501.07806cs.CV2025-01中稿 · IEEE Transactions …被引 15

融合运动与时间线索,提升无监督视频目标分割精度

Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation

  • 在编码器中联合提取外观与运动特征,增强表征互补性
  • 引入时序变换模块,捕捉长距离帧间动态关系
  • 多级解码器协同优化,实现精准且鲁棒的物体定位

本文针对无监督视频对象分割(UVOS)的挑战,提出一种高效算法MTNet,同时利用运动与时间线索。不同于以往仅关注外观与运动融合或建模时间关系的方法,本方法在统一框架内整合两者。通过在编码器中有效融合外观与运动特征,促进更互补的表示;引入时序变换模块以捕捉视频中复杂的长程上下文动态;采用多级解码器架构,充分挖掘各层级特征,生成越来越精确的分割掩码。实验表明,MTNet在多个基准上均达到最先进的无监督视频对象分割性能,并在视频显著性目标检测任务中表现竞争力,展现了方法在多种分割任务中的鲁棒性与适应性。代码已开源。

原文摘要 · Abstract (English)

In this paper, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues. Unlike previous methods that focus solely on integrating appearance with motion or on modeling temporal relations, our method combines both aspects by integrating them within a unified framework. MTNet is devised by effectively merging appearance and motion features during the feature extraction process within encoders, promoting a more complementary representation. To capture the intricate long-range contextual dynamics and information embedded within videos, a temporal transformer module is introduced, facilitating efficacious inter-frame interactions throughout a video clip. Furthermore, we employ a cascade of decoders all feature levels across all feature levels to optimally exploit the derived features, aiming to generate increasingly precise segmentation masks. As a result, MTNet provides a strong and compact framework that explores both temporal and cross-modality knowledge to robustly localize and track the primary object accurately in various challenging scenarios efficiently. Extensive experiments across diverse benchmarks conclusively show that our method not only attains state-of-the-art performance in unsupervised video object segmentation but also delivers competitive results in video salient object detection. These findings highlight the method's robust versatility and its adeptness in adapting to a range of segmentation tasks. Source code is available on https://github.com/hy0523/MTNet.

视频分割无监督学习时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。