arXiv:2508.10731cs.CVcs.LG2025-08ICCV被引 13

通过分解物体为视觉基元,实现对新类别更可靠的发现。

Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction

  • 将物体拆解为视觉基元,用语义重建构建表征。
  • 融合主导与上下文共识,提升对新类别的识别能力。
  • 适合研究新类别发现和人类认知启发的模型设计者。

人类感知系统在已知和未知类别间识别物体的能力远超当前机器学习框架。尽管广义类别发现(GCD)旨在弥合这一差距,现有方法主要聚焦于优化目标函数。本文提出一种正交解决方案,受人类理解新物体的认知过程启发:将物体分解为视觉基元,并建立跨知识的比较。我们提出ConGCD,通过高层语义重建建立基元导向的表征,利用去构造机制绑定类内共享属性。模仿人类在视觉处理中的偏好多样性——不同个体依赖主导或上下文线索——我们设计了主导共识单元与上下文共识单元,分别捕捉类别判别性模式与固有分布不变性。共识调度器动态优化激活路径,最终通过多重共识集成生成预测。在粗粒度与细粒度基准上的广泛评估表明,ConGCD作为共识感知范式具有显著有效性。代码已公开于github.com/lytang63/ConGCD。

原文摘要 · Abstract (English)

Human perceptual systems excel at inducing and recognizing objects across both known and novel categories, a capability far beyond current machine learning frameworks. While generalized category discovery (GCD) aims to bridge this gap, existing methods predominantly focus on optimizing objective functions. We present an orthogonal solution, inspired by the human cognitive process for novel object understanding: decomposing objects into visual primitives and establishing cross-knowledge comparisons. We propose ConGCD, which establishes primitive-oriented representations through high-level semantic reconstruction, binding intra-class shared attributes via deconstruction. Mirroring human preference diversity in visual processing, where distinct individuals leverage dominant or contextual cues, we implement dominant and contextual consensus units to capture class-discriminative patterns and inherent distributional invariants, respectively. A consensus scheduler dynamically optimizes activation pathways, with final predictions emerging through multiplex consensus integration. Extensive evaluations across coarse- and fine-grained benchmarks demonstrate ConGCD's effectiveness as a consensus-aware paradigm. Code is available at github.com/lytang63/ConGCD.

类别发现认知启发共识机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。