arXiv:2602.19910cs.CV2026-02中稿 · CVPR

通过自监督率减少学习多模态表示,提升未知类别发现能力

Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery

  • 引入半监督率减少框架,强化模态内关系对齐
  • 在通用与细粒度数据集上优于现有方法
  • 适合需要开放集识别的多模态场景

广义类别发现(GCD)旨在识别已知和未知类别,仅提供部分已知类别的标签,属于具有挑战性的开集识别问题。当前主流方法依赖多模态表示学习,高度依赖模态间对齐,但缺乏对模态内对齐的有效建模以形成理想的表示分布结构。本文提出一种基于半监督率减少的新型多模态表示学习框架SSR²-GCD,通过强调模态内关系对齐,学习具备理想结构特性的跨模态表示。为进一步促进知识迁移,利用视觉语言模型提供的模态间对齐整合提示候选。在通用与细粒度基准数据集上进行了广泛实验,验证了该方法的优越性能。

原文摘要 · Abstract (English)

Generalized Category Discovery (GCD) aims to identify both known and unknown categories, with only partial labels given for the known categories, posing a challenging open-set recognition problem. State-of-the-art approaches for GCD task are usually built on multi-modality representation learning, which is heavily dependent upon inter-modality alignment. However, few of them cast a proper intra-modality alignment to generate a desired underlying structure of representation distributions. In this paper, we propose a novel and effective multi-modal representation learning framework for GCD via Semi-Supervised Rate Reduction, called SSR$^2$-GCD, to learn cross-modality representations with desired structural properties based on emphasizing to properly align intra-modality relationships. Moreover, to boost knowledge transfer, we integrate prompt candidates by leveraging the inter-modal alignment offered by Vision Language Models. We conduct extensive experiments on generic and fine-grained benchmark datasets demonstrating superior performance of our approach.

多模态类别发现自监督开放集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。