无需人工标注,直接在场景图像上实现高质量无监督全景分割。
Scene-Centric Unsupervised Panoptic Segmentation
- 基于视觉、深度和运动信息生成高分辨率伪标签
- 在Cityscapes上达到9.4%的全景质量提升,超越现有最优方法
- 适合无标注数据场景下的全景分割研究与应用
无监督全景分割旨在不依赖人工标注数据的情况下,将图像划分为语义有意义的区域和独立的对象实例。与以往基于对象中心数据的无监督全景理解方法不同,本文摒弃了对对象中心训练数据的需求,实现了对复杂场景的无监督理解。为此,我们提出了首个直接在场景中心图像上训练的无监督全景分割方法。具体而言,我们提出一种方法,在复杂场景图像上结合视觉表征、深度和运动线索,生成高分辨率的全景伪标签。通过伪标签训练与全景自训练策略相结合,该方法在无需任何人工标注的前提下,准确预测复杂场景的全景分割结果。实验表明,该方法显著提升了全景分割质量,例如在Cityscapes数据集上相比近期最优方法提升了9.4个百分点的全景质量(PQ)。
原文摘要 · Abstract (English)
Unsupervised panoptic segmentation aims to partition an image into semantically meaningful regions and distinct object instances without training on manually annotated data. In contrast to prior work on unsupervised panoptic scene understanding, we eliminate the need for object-centric training data, enabling the unsupervised understanding of complex scenes. To that end, we present the first unsupervised panoptic method that directly trains on scene-centric imagery. In particular, we propose an approach to obtain high-resolution panoptic pseudo labels on complex scene-centric data, combining visual representations, depth, and motion cues. Utilizing both pseudo-label training and a panoptic self-training strategy yields a novel approach that accurately predicts panoptic segmentation of complex scenes without requiring any human annotations. Our approach significantly improves panoptic quality, e.g., surpassing the recent state of the art in unsupervised panoptic segmentation on Cityscapes by 9.4% points in PQ.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。