构建多粒度视频目标分割数据集,提升模型对非显著目标的追踪能力。
Multi-Granularity Video Object Segmentation
- 构建包含显著与非显著目标的多粒度标注数据集
- 提出内存式掩码传播模型,在新数据集上性能最优
- 适合关注真实场景下细粒度目标分割的研究者
当前视频分割基准仅标注显著对象(即前景实例),尽管以往方法架构精巧,但在真实场景中仍难以适应。因此,亟需构建面向视频场景中多粒度分割目标的新型数据集。本文提出大规模、密集标注的多粒度视频对象分割(MUG-VOS)数据集,涵盖多种类型和粒度的掩码标注。我们自动构建了训练集以支持显著与非显著对象的跟踪,并人工标注了测试集用于可靠评估。此外,基于MUG-VOS提出了内存式掩码传播模型(MMPM),在现有视频对象分割方法及Segment SAM基础上均取得最佳性能。项目页面见https://cvlab-kaist.github.io/MUG-VOS。
原文摘要 · Abstract (English)
Current benchmarks for video segmentation are limited to annotating only salient objects (i.e., foreground instances). Despite their impressive architectural designs, previous works trained on these benchmarks have struggled to adapt to real-world scenarios. Thus, developing a new video segmentation dataset aimed at tracking multi-granularity segmentation target in the video scene is necessary. In this work, we aim to generate multi-granularity video segmentation dataset that is annotated for both salient and non-salient masks. To achieve this, we propose a large-scale, densely annotated multi-granularity video object segmentation (MUG-VOS) dataset that includes various types and granularities of mask annotations. We automatically collected a training set that assists in tracking both salient and non-salient objects, and we also curated a human-annotated test set for reliable evaluation. In addition, we present memory-based mask propagation model (MMPM), trained and evaluated on MUG-VOS dataset, which leads to the best performance among the existing video object segmentation methods and Segment SAM-based video segmentation methods. Project page is available at https://cvlab-kaist.github.io/MUG-VOS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。