用视觉代理提升组合零样本学习的识别能力
Learning Visual Proxy for Compositional Zero-Shot Learning
- 为属性、物体及组合初始化视觉代理,优化视觉空间捕捉细粒度特征
- 跨模态联合学习使未见组合识别准确率提升,相似对区分更精准
- 适合研究零样本学习、视觉语言模型融合的学者参考
组合零样本学习(CZSL)旨在通过已知组合的知识识别新颖的属性-对象组合。现有方法利用视觉-语言模型(VLMs)对齐文本原型与视觉特征,但存在两大局限:(1) 模态差距阻碍语义相似对的区分;(2) 单模态文本原型缺乏细粒度视觉线索。本文提出视觉代理学习(Visual Proxy Learning),通过文本表示初始化属性、物体及其组合的视觉代理,并优化视觉空间以捕获细粒度特征,增强视觉表征。此外,提出跨模态联合学习(CMJL),在文本-图像空间与细粒度视觉空间间施加跨模态约束,提升未见组合的泛化能力并更好区分相似组合。在四个CZSL基准上,闭世界场景下达到领先性能,开世界设置下表现竞争力,验证了方法在组合泛化上的有效性。
原文摘要 · Abstract (English)
Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Current methods align textual prototypes with visual features via Vision-Language Models (VLMs), but suffer from two limitations: (1) modality gaps hinder the discrimination of semantically similar pairs, and (2) single-modal textual prototypes lack fine-grained visual cues. In this paper, we introduce Visual Proxy Learning, a method that reduces modality gaps and enhances compositional generalization. We initialize visual proxies for attributes, objects, and their compositions using text representations and optimize the visual space to capture fine-grained cues, improving visual representations. Additionally, we propose Cross-Modal Joint Learning (CMJL), which imposes cross-modal constraints between the text-image and fine-grained visual spaces, improving generalization for unseen compositions and discriminating similar pairs. Experiments show state-of-the-art performance in closed-world scenarios and competitive results in open-world settings across four CZSL benchmarks, demonstrating the effectiveness of our approach in compositional generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。