通过视觉与语义提示协同,提升零样本图像识别准确率
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
- 用视觉提示和语义提示分别捕捉特征与语义关联
- 浅层弱融合、深层强融合,实现提示信息有效整合
- 在多个零样本学习基准上表现优于现有方法
通用零样本学习旨在利用类间共享的语义信息,识别已见和未见类别。这要求视觉与语义保持一致对齐。现有方法通过已见类别数据微调视觉主干网络以获得相关语义的视觉特征,但在训练图像有限时易对已见类别过拟合。本文提出一种新型视觉与语义提示协同框架,采用提示调优技术实现高效特征适配。具体地,设计视觉提示以整合判别性视觉信息,设计语义提示以整合语义信息以实现视觉-语义对齐。为实现有效提示融合,进一步在浅层设计弱融合机制,在深层设计强融合机制。通过视觉与语义提示的协作,可获得具有判别性的语义相关特征,用于通用零样本图像识别。大量实验表明,该框架在传统零样本学习与通用零样本学习基准上均显著优于其他先进方法。
原文摘要 · Abstract (English)
Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires consistent visual-semantic alignment. Existing approaches fine-tune the visual backbone by seen-class data to obtain semantic-related visual features, which may cause overfitting on seen classes with a limited number of training images. This paper proposes a novel visual and semantic prompt collaboration framework, which utilizes prompt tuning techniques for efficient feature adaptation. Specifically, we design a visual prompt to integrate the visual information for discriminative feature learning and a semantic prompt to integrate the semantic formation for visualsemantic alignment. To achieve effective prompt information integration, we further design a weak prompt fusion mechanism for the shallow layers and a strong prompt fusion mechanism for the deep layers in the network. Through the collaboration of visual and semantic prompts, we can obtain discriminative semantic-related features for generalized zero-shot image recognition. Extensive experiments demonstrate that our framework consistently achieves favorable performance in both conventional zero-shot learning and generalized zero-shot learning benchmarks compared to other state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。