arXiv:2409.11624cs.CVcs.LG2024-09CVPR被引 9

解决多模态数据中的新类别发现难题,提升开放世界识别能力

Multimodal Generalized Category Discovery

  • 通过对比学习与知识蒸馏对齐多模态特征与输出空间
  • 在UPMC-Food101和N24News上分别提升11.5%和4.7%性能
  • 适合处理图文、音视频等多源信息的开放世界场景

广义类别发现(GCD)旨在将输入分类为已知与新类别,对开放世界科学发现至关重要。然而现有方法仅限于单模态数据,忽视了现实数据固有的多模态特性。本文将GCD拓展至多模态场景,利用不同模态间互补信息提升性能。理论分析与实验证明,多模态GCD的核心挑战在于跨模态异构信息的有效对齐。为此提出MM-GCD框架,通过对比学习与知识蒸馏技术对齐各模态的特征空间与输出空间。在UPMC-Food101和N24News数据集上,该方法分别超越先前最佳方法11.5%和4.7%,达到新基准性能。

原文摘要 · Abstract (English)

Generalized Category Discovery (GCD) aims to classify inputs into both known and novel categories, a task crucial for open-world scientific discoveries. However, current GCD methods are limited to unimodal data, overlooking the inherently multimodal nature of most real-world data. In this work, we extend GCD to a multimodal setting, where inputs from different modalities provide richer and complementary information. Through theoretical analysis and empirical validation, we identify that the key challenge in multimodal GCD lies in effectively aligning heterogeneous information across modalities. To address this, we propose MM-GCD, a novel framework that aligns both the feature and output spaces of different modalities using contrastive learning and distillation techniques. MM-GCD achieves new state-of-the-art performance on the UPMC-Food101 and N24News datasets, surpassing previous methods by 11.5\% and 4.7\%, respectively.

多模态类别发现对比学习开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。