首个无监督视频全景分割方法,靠场景线索自动生成标签。
Scene-Centric Unsupervised Video Panoptic Segmentation

- 基于深度、运动和视觉线索生成时序一致的伪标签
- 在Cityscapes上达到34.7% mPS,优于所有基线
- 适合对标注成本敏感的视频理解研究
视频全景分割(VPS)旨在联合检测、分割并追踪所有物体,同时将视频划分为语义一致的区域。本文提出无监督视频全景分割任务,无需人工标注。现有无监督场景理解工作多集中于图像分割,视频领域仍待探索。我们提出VideoCUPS,首个无监督VPS方法,通过利用无监督深度、运动和视觉线索,从场景中心视频中生成时序一致的全景视频伪标签。使用新型Video DropLoss在这些伪标签上训练,获得高精度的无监督VPS模型。为评估进展,我们构建了全面评测协议及四个竞争性基线,扩展了当前最先进的无监督全景图像与实例视频分割模型至VPS。VideoCUPS超越所有基线,在标签效率上表现优异。本工作提供了无监督VPS研究的坚实基础。
原文摘要 · Abstract (English)
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focused on image segmentation tasks; the video domain remains underexplored. We propose VideoCUPS, the first unsupervised VPS approach. VideoCUPS generates temporally consistent panoptic video pseudo-labels from scene-centric videos by exploiting unsupervised depth, motion, and visual cues. Training on these pseudo-labels using a novel Video DropLoss yields an accurate, unsupervised VPS model. To benchmark progress, we introduce a comprehensive evaluation protocol and four competitive baselines, extending state-of-the-art unsupervised panoptic image and instance video segmentation models to VPS. VideoCUPS outperforms all baselines and demonstrates strong label-efficient learning. With VideoCUPS, our evaluation protocol, and baselines, we provide a strong foundation for future research on unsupervised VPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。