arXiv:2507.22061cs.CV2025-07ICCV被引 6

构建新数据集MOVE,解决视频目标分割中运动模式理解难题

MOVE: Motion-Guided Few-Shot Video Object Segmentation

  • 基于运动模式设计新数据集,突破传统静态类别限制
  • 现有方法在运动引导分割任务上表现不佳,平均mIoU不足30%
  • 提出解耦运动与外观的基准模型,显著提升少样本运动理解能力

本文针对运动引导的少样本视频目标分割(FSVOS)问题,旨在仅凭少量标注样本和相同运动模式实现动态物体分割。现有数据集与方法多关注静态类别属性,忽视视频中的丰富时序动态,限制了其在需运动理解场景的应用。为此,我们构建了大规模的MOVE数据集,专门用于运动引导的少样本视频目标分割。基于MOVE,我们在两个实验设置下全面评估了来自三个相关任务的6种先进方法。结果表明,当前方法在运动引导分割任务上表现不佳,由此分析挑战并提出一种基准方法——解耦运动与外观网络(DMA)。实验显示,该方法在少样本运动理解方面表现优异,为后续研究奠定了坚实基础。

原文摘要 · Abstract (English)

This work addresses motion-guided few-shot video object segmentation (FSVOS), which aims to segment dynamic objects in videos based on a few annotated examples with the same motion patterns. Existing FSVOS datasets and methods typically focus on object categories, which are static attributes that ignore the rich temporal dynamics in videos, limiting their application in scenarios requiring motion understanding. To fill this gap, we introduce MOVE, a large-scale dataset specifically designed for motion-guided FSVOS. Based on MOVE, we comprehensively evaluate 6 state-of-the-art methods from 3 different related tasks across 2 experimental settings. Our results reveal that current methods struggle to address motion-guided FSVOS, prompting us to analyze the associated challenges and propose a baseline method, Decoupled Motion Appearance Network (DMA). Experiments demonstrate that our approach achieves superior performance in few shot motion understanding, establishing a solid foundation for future research in this direction.

视频分割少样本学习运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。