arXiv:2507.14935cs.CV2025-07ICCV被引 5

解决跨模态模型在开放场景下的泛化难题,提升对未知类别的识别能力。

Open-set Cross Modal Generalization via Multimodal Unified Representation

  • 提出细粒度掩码对比学习与跨模态拼图任务,增强多模态对齐与特征多样性
  • 在新提出OSCMG任务上实现9.2%的性能提升,显著优于现有方法
  • 适合关注开放世界多模态学习、真实场景泛化的研究者

本文将跨模态泛化(CMG)拓展至开放集环境,提出更具挑战性的开放集跨模态泛化(OSCMG)任务。该任务评估多模态统一表征在开放集条件下的表现,弥补了以往封闭集评价的不足。OSCMG不仅要求跨模态知识迁移,还需在新模态中对未见类别具备鲁棒泛化能力,这在真实应用中极为常见。现有工作未考虑开放集场景。为此,我们提出MICU框架,包含两个关键组件:细粒度掩码多模态InfoNCE(FCMI)和跨模态统一拼图(CUJP)。FCMI通过在整体语义与时间层级上进行对比学习并引入掩码机制,提升对齐与泛化能力;CUJP结合模态无关特征选择与自监督学习,增强特征多样性与模型不确定性建模,从而提升处理未知类别的能力。在CMG与新提出的OSCMG任务上的大量实验验证了方法的有效性。代码已开源。

原文摘要 · Abstract (English)

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in open-set conditions, addressing the limitations of prior closed-set cross-modal evaluations. OSCMG requires not only cross-modal knowledge transfer but also robust generalization to unseen classes within new modalities, a scenario frequently encountered in real-world applications. Existing multimodal unified representation work lacks consideration for open-set environments. To tackle this, we propose MICU, comprising two key components: Fine-Coarse Masked multimodal InfoNCE (FCMI) and Cross modal Unified Jigsaw Puzzles (CUJP). FCMI enhances multimodal alignment by applying contrastive learning at both holistic semantic and temporal levels, incorporating masking to enhance generalization. CUJP enhances feature diversity and model uncertainty by integrating modality-agnostic feature selection with self-supervised learning, thereby strengthening the model's ability to handle unknown categories in open-set tasks. Extensive experiments on CMG and the newly proposed OSCMG validate the effectiveness of our approach. The code is available at https://github.com/haihuangcode/CMG.

跨模态开放集统一表征自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。