arXiv:2501.13859cs.CV2025-01ICCV被引 2

用视觉代理提升组合零样本学习的识别能力

Learning Visual Proxy for Compositional Zero-Shot Learning

  • 为属性、物体及组合初始化视觉代理,优化视觉空间捕捉细粒度特征
  • 跨模态联合学习使未见组合识别准确率提升,相似对区分更精准
  • 适合研究零样本学习、视觉语言模型融合的学者参考

组合零样本学习(CZSL)旨在通过已知组合的知识识别新颖的属性-对象组合。现有方法利用视觉-语言模型(VLMs)对齐文本原型与视觉特征,但存在两大局限:(1) 模态差距阻碍语义相似对的区分;(2) 单模态文本原型缺乏细粒度视觉线索。本文提出视觉代理学习(Visual Proxy Learning),通过文本表示初始化属性、物体及其组合的视觉代理,并优化视觉空间以捕获细粒度特征,增强视觉表征。此外,提出跨模态联合学习(CMJL),在文本-图像空间与细粒度视觉空间间施加跨模态约束,提升未见组合的泛化能力并更好区分相似组合。在四个CZSL基准上,闭世界场景下达到领先性能,开世界设置下表现竞争力,验证了方法在组合泛化上的有效性。

原文摘要 · Abstract (English)

Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Current methods align textual prototypes with visual features via Vision-Language Models (VLMs), but suffer from two limitations: (1) modality gaps hinder the discrimination of semantically similar pairs, and (2) single-modal textual prototypes lack fine-grained visual cues. In this paper, we introduce Visual Proxy Learning, a method that reduces modality gaps and enhances compositional generalization. We initialize visual proxies for attributes, objects, and their compositions using text representations and optimize the visual space to capture fine-grained cues, improving visual representations. Additionally, we propose Cross-Modal Joint Learning (CMJL), which imposes cross-modal constraints between the text-image and fine-grained visual spaces, improving generalization for unseen compositions and discriminating similar pairs. Experiments show state-of-the-art performance in closed-world scenarios and competitive results in open-world settings across four CZSL benchmarks, demonstrating the effectiveness of our approach in compositional generalization.

零样本学习视觉代理跨模态学习组合识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。