arXiv:2412.06190cs.CV2024-12中稿 · IEEE Transactions …被引 2

针对开放词汇多标签识别中类别间语义关联捕捉不足的问题,提出自适应跨模态语义优化与迁移框架。

Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition

  • 根据目标类别自适应选择关键视觉区域,提升语义表征精度
  • 构建类别自适应相关图,将已知类别知识迁移至未知类别
  • 适用于开放词汇场景下的多标签图像识别任务

得益于CLIP的泛化能力,近期的视觉语言预训练(VLP)模型在日常图像中捕获广泛视觉概念方面表现出色。然而,在开放词汇设置下存在未见类别,现有算法难以捕捉类别间的语义关联,导致开放词汇多标签识别(OV-MLR)性能不佳。此外,不同物体类别间判别区域数量差异显著,而现有方法采用固定数量的图像块匹配,引入噪声视觉线索,阻碍目标语义的准确提取。为此,本文提出一种新的类别自适应跨模态语义精炼与迁移(C²SRT)框架,以类别自适应方式建模类别内部及类别间的语义关联。该框架包含两个互补模块:类别内语义精炼(ISR)模块利用VLP模型的跨模态知识,自适应选择代表目标类别的局部判别区域;类别间语义迁移(IST)模块通过构建类别自适应相关图,自适应发现目标类别的相关类别,并将已知类别语义知识迁移至未见类别。在多个OV-MLR基准上的实验表明,所提框架优于现有方法。

原文摘要 · Abstract (English)

Benefiting from the generalization capability of CLIP, recent vision language pre-training (VLP) models have demonstrated the ability to capture a wide range of visual concepts in daily images. However, due to the presence of unseen categories in open-vocabulary settings, existing algorithms struggle to capture semantic correlations between categories, leading to suboptimal performance on open-vocabulary multi-label recognition (OV-MLR). Furthermore, the substantial variation in the number of discriminative areas across diverse object categories is misaligned with the fixed-number patch matching used in current methods, introducing noisy visual cues that hinder the capture of target semantics. To address these challenges, we propose a novel category-adaptive cross-modal semantic refinement and transfer (C$^2$SRT) framework to model semantic correlations both within each category and across different categories, in a category-adaptive manner. The proposed framework consists of two complementary modules, i.e., intra-category semantic refinement (ISR) module and inter-category semantic transfer (IST) module. Specifically, the ISR module leverages the cross-modal knowledge of the VLP model to adaptively select a set of local discriminative regions that represent the semantics of the target category. The IST module adaptively discovers a set of correlated categories for a target category by constructing a category-adaptive correlation graph and transfers semantic knowledge from the correlated seen categories to unseen ones. Experiments on OV-MLR benchmarks demonstrate that the proposed C$^2$SRT framework improves over current methods.

开放词汇多标签识别跨模态语义迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。