arXiv:2505.02071cs.CVcs.LG2025-05CVPR被引 1

用紧凑性聚类机制实现无监督对象发现,自动生成高质量分割掩码。

Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning

  • 基于紧凑性原理的注意力聚类模块,构建分层网络提取对象中心表征。
  • 在六大数据集上优于或媲美现有模型,九项指标表现优异。
  • 无需预设对象数量,能更好分割背景,适合无监督图像分割任务。

我们提出紧凑聚类注意力(COCA)层,一种有效的构建模块,通过分层策略实现对象中心表征学习,并解决单图上的无监督对象发现任务。当级联到自下而上的分层网络架构(即COCA-Net)时,COCA是一种基于注意力的聚类模块,能够从多对象场景中提取对象中心表征。其核心是利用物理紧凑性概念的新型聚类算法,突出场景中的独立对象质心,提供空间归纳偏置。得益于该策略,COCA-Net在解码器侧和编码器侧均生成高质量分割掩码。此外,COCA-Net不依赖预设的对象掩码数量,对背景元素的分割效果优于竞争对手。我们在六个广泛采用的数据集上展示了COCA-Net的分割性能,其在九个不同评估指标上达到优于或可比当前最优模型的结果。

原文摘要 · Abstract (English)

We propose the Compact Clustering Attention (COCA) layer, an effective building block that introduces a hierarchical strategy for object-centric representation learning, while solving the unsupervised object discovery task on single images. COCA is an attention-based clustering module capable of extracting object-centric representations from multi-object scenes, when cascaded into a bottom-up hierarchical network architecture, referred to as COCA-Net. At its core, COCA utilizes a novel clustering algorithm that leverages the physical concept of compactness, to highlight distinct object centroids in a scene, providing a spatial inductive bias. Thanks to this strategy, COCA-Net generates high-quality segmentation masks on both the decoder side and, notably, the encoder side of its pipeline. Additionally, COCA-Net is not bound by a predetermined number of object masks that it generates and handles the segmentation of background elements better than its competitors. We demonstrate COCA-Net's segmentation performance on six widely adopted datasets, achieving superior or competitive results against the state-of-the-art models across nine different evaluation metrics.

无监督学习对象发现注意力机制图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。