通过自适应部件学习提升细粒度分类发现的准确性与泛化能力
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
- 用可学习部件查询和DINO先验生成一致物体部件及其对应关系
- 设计全最小对比损失,兼顾判别性与泛化性,显著提升识别效果
- 无需额外标注,可即插即用,适用于多种细粒度GCD框架
广义类别发现(GCD)旨在通过区分新旧类别,识别已知和未知类别的未标注图像,并从另一组已知类别的标记图像中迁移知识。现有GCD方法依赖于如DINO等自监督视觉变换器进行表征学习。然而,仅关注DINO CLS标记的全局表征会带来判别性与泛化性之间的固有权衡。本文提出自适应部件发现与学习方法APL,利用一组共享的可学习部件查询和DINO部件先验,在不同相似图像间生成一致的物体部件及其对应关系,无需额外标注。更重要的是,我们提出一种新颖的全最小对比损失,以学习具有判别性且泛化性强的部件表征:自适应突出有助于区分相似类别的关键部件,增强判别性;同时共享其他部件以促进知识迁移,提升泛化能力。APL可通过替换现有GCD框架中的CLS标记特征,轻松集成到不同框架中,在细粒度数据集上表现出显著提升。
原文摘要 · Abstract (English)
Generalized Category Discovery (GCD) aims to recognize unlabeled images from known and novel classes by distinguishing novel classes from known ones, while also transferring knowledge from another set of labeled images with known classes. Existing GCD methods rely on self-supervised vision transformers such as DINO for representation learning. However, focusing solely on the global representation of the DINO CLS token introduces an inherent trade-off between discriminability and generalization. In this paper, we introduce an adaptive part discovery and learning method, called APL, which generates consistent object parts and their correspondences across different similar images using a set of shared learnable part queries and DINO part priors, without requiring any additional annotations. More importantly, we propose a novel all-min contrastive loss to learn discriminative yet generalizable part representation, which adaptively highlights discriminative object parts to distinguish similar categories for enhanced discriminability while simultaneously sharing other parts to facilitate knowledge transfer for improved generalization. Our APL can easily be incorporated into different GCD frameworks by replacing their CLS token feature with our part representations, showing significant enhancements on fine-grained datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。