arXiv:2606.21446cs.CV2026-06

提升多模态分类中图文协同,更好发现新类别

Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery

论文配图:Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery
图 1 · 摘自论文原文
  • 在编码阶段注入视觉信息增强文本特征学习
  • 通过局部邻域一致性提升新旧类别的识别精度
  • 可插拔适配现有方法,适合多模态发现任务

广义类别发现(GCD)旨在对已知类别进行分类并从无标签数据中发现新类别。近期多模态方法在双分支架构中引入检索或合成文本,为视觉特征提供语义补充。然而现有方法的跨模态协同仍较粗糙:两模态独立编码,文本编码时未消除偏差与噪声,且现有互学习策略仅基于全局类别锚点,缺乏细粒度关系监督。为此,我们提出协同双分支自适应框架(SDBA),可作为插件增强已有双分支方法(如GET、TextGCD)。SDBA包含两个组件:跨模态协同适配器在编码层向文本分支注入视觉信息,提升文本特征学习;邻域互学习模块通过双向KL散度强制两分支局部邻域分布一致,为新旧类别提供细粒度关系监督。六组基准测试显示,SDBA达到当前最优性能,不同基线上的持续提升验证了其广泛适用性。

原文摘要 · Abstract (English)

Generalized Category Discovery (GCD) aims to classify old categories and discover new ones from unlabeled data. Recent multi-modal approaches introduce retrieved or synthesized texts into a dual-branch architecture to provide semantic cues complementary to visual features. However, the cross-modal synergy in existing dual-branch methods remains coarse and incomplete: the two modalities are encoded independently with the bias and noise in the derived text left unaddressed during encoding, and existing mutual learning strategies operate only on global class-level anchors, lacking fine-grained relational supervision. To address these limitations, we propose the Synergistic Dual-Branch Adaptation (SDBA) framework, which serves as a plug-and-play enhancement compatible with existing dual-branch methods such as GET and TextGCD. SDBA comprises two components: the cross-modal synergistic adapter inserts lightweight adapters into both branches and further injects visual information into the text adapter at each encoder layer to enhance text feature learning during encoding; the neighborhood mutual learning module enforces consistent local neighborhood distributions between the two branches via bidirectional KL divergence, providing fine-grained relational supervision for both old and new classes. Extensive experiments on six benchmarks demonstrate state-of-the-art performance, and consistent improvements on different baselines validate the broad scalability of the proposed framework.

多模态类别发现双分支协同学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。