用先验约束关联学习,提升细粒度类别发现效果。
Prior-Constrained Association Learning for Fine-Grained Generalized Category Discovery
- 引入已知类别的先验信息,约束无标签数据的关联过程。
- 在多个基准上优于现有方法,显著提升类别发现准确率。
- 适合需要从未知类别中挖掘细粒度语义的研究者。
本文针对广义类别发现(GCD)任务,即在已知类别标注样本的帮助下,对可能包含新类别的无标签数据进行聚类。与传统半监督学习相比,GCD更具挑战性,因无标签数据可能来自未见的新类别。当前先进方法通常依赖自蒸馏的参数化分类器,但未能充分利用实例间的相似性来发现类别特异性语义,而这对于表征学习和类别发现至关重要。为此,本文重新审视基于关联的范式,提出先验约束关联学习(Prior-constrained Association Learning),通过已知类别的标注数据提供唯一先验,指导无标签数据的语义关联。不同于以往仅将先验用于预或后处理聚类,本方法将其深度融入关联过程,引导生成可靠分组。利用非参数原型对比增强表征学习,并结合参数化与非参数化分类,互补提升性能。在多个GCD基准上进行了充分实验,验证了所提方法的有效性。
原文摘要 · Abstract (English)
This paper addresses generalized category discovery (GCD), the task of clustering unlabeled data from potentially known or unknown categories with the help of labeled instances from each known category. Compared to traditional semi-supervised learning, GCD is more challenging because unlabeled data could be from novel categories not appearing in labeled data. Current state-of-the-art methods typically learn a parametric classifier assisted by self-distillation. While being effective, these methods do not make use of cross-instance similarity to discover class-specific semantics which are essential for representation learning and category discovery. In this paper, we revisit the association-based paradigm and propose a Prior-constrained Association Learning method to capture and learn the semantic relations within data. In particular, the labeled data from known categories provides a unique prior for the association of unlabeled data. Unlike previous methods that only adopts the prior as a pre or post-clustering refinement, we fully incorporate the prior into the association process, and let it constrain the association towards a reliable grouping outcome. The estimated semantic groups are utilized through non-parametric prototypical contrast to enhance the representation learning. A further combination of both parametric and non-parametric classification complements each other and leads to a model that outperforms existing methods by a significant margin. On multiple GCD benchmarks, we perform extensive experiments and validate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。