arXiv:2607.00620cs.CVcs.AI2026-07中稿 · ICML

通过分解图像为可复用的视觉原语,提升开放世界中新类别发现能力。

Identifying Latent Concepts and Structures for Generalized Category Discovery

论文配图:Identifying Latent Concepts and Structures for Generalized Category Discovery
图 1 · 摘自论文原文
  • 用低秩原语混合重构特征,打破高秩纠缠表示
  • 在多个基线模型上提升新类别识别准确率,最高增益达8.2%
  • 适合需要强泛化能力的开放世界场景应用

广义类别发现(GCD)旨在开放世界中识别已知类别并自主发现新类别。然而,现有方法多聚焦于聚类目标设计,忽视了关键瓶颈:标准视觉主干网络产生高秩、纠缠的标记表示,不利于无监督发现潜在概念与结构。本文提出组合原语场(CPF-GCD),一种新型表示学习框架,通过强制低秩组合组织重塑特征空间,使潜在结构可识别。核心假设是所有类别(已知或新)均可由有限可学习的视觉原语及其空间排列构成。CPF通过空间场机制实现此几何约束,插入主干与分类头之间,以低秩原语混合重写噪声补丁标记,有效将图像分解为可复用的基本单元及其空间布局。通过显式建模原语的空间分布,新类别自然表现为共享词汇上的新激活模式。这使得表征重点从全局嵌入划分转向构建结构化可分离的原语场。大量实验表明,CPF作为通用即插即用模块,在多种GCD基线上持续提升性能,验证了识别与利用低秩组合结构是开放世界识别的关键归纳偏置。

原文摘要 · Abstract (English)

Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. However, current approaches primarily focus on designing clustering objectives, often overlooking a critical bottleneck: standard vision backbones yield high-rank, entangled token representations that are ill-suited for unsupervised discovery of latent concepts and structures. In this paper, we propose Compositional Primitive Fields (CPF-GCD), a novel representation learning framework that reshapes the feature space to make such latent structure identifiable by enforcing a low-rank compositional organization. Our core hypothesis is that all categories, whether known or novel, can be expressed as compositions and spatial arrangements of a finite set of learnable visual primitives that capture reusable concepts. CPF instantiates this geometric constraint via a spatial field mechanism. Inserted between the backbone and the head, it rewrites noisy patch tokens through low-rank primitive mixtures, effectively decomposing images into reusable atomic parts and their spatial layouts. By explicitly modeling the spatial distribution of primitives, CPF enables novel categories to emerge naturally as new activation patterns over a shared vocabulary. This shifts the focus of representation from merely partitioning global embeddings to constructing a structured and separable primitive field. Extensive experiments demonstrate that CPF serves as a generic, plug-and-play module that consistently boosts performance across diverse GCD baselines, validating that identifying and leveraging low-rank compositional structure is a crucial inductive bias for open-world recognition.

类别发现视觉原语开放世界表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。