arXiv:2602.23114cs.CV2026-02

测试时动态积累多模态知识,提升组合零样本学习的泛化能力

WARM-CAT: Warm-Started Test-Time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning

  • 测试时从无监督数据中更新图文双模态原型,缓解标签分布偏移
  • 引入自适应权重与动态优先队列,实现高置信图像的持续学习
  • 新基准C-Fashion和优化后的MIT-States,推动领域可靠评估

组合零样本学习(CZSL)旨在基于已见组合的知识识别未见的属性-对象组合。现有方法因测试时标签空间分布偏移而性能下降,该偏移源于属性与对象重新组合产生的未见组合。为此,本文提出一种新方法,在测试时从无监督数据中累积文本与视觉模态的综合知识,更新多模态原型。在此基础上,设计自适应更新权重以控制原型调整程度,使模型灵活应对测试时的分布偏移。此外,引入动态优先队列存储高置信图像,利用历史图像获取视觉原型用于推理。由于模型在测试时倾向于依赖队列中已有组合,本文通过初始化队列并利用已见与未见文本原型间的映射生成未见视觉原型,实现预热。考虑多模态知识语义一致性,通过多模态协同表示学习对齐文本与视觉原型。为提供更可靠的CZSL评估,引入新基准数据集C-Fashion,同时清理广泛使用的有噪声MIT-States数据集。大量实验表明,该方法在四个基准数据集上,于闭世界与开世界设置下均达到当前最优性能。

原文摘要 · Abstract (English)

Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions based on the knowledge learned from seen ones. Existing methods suffer from performance degradation caused by the distribution shift of label space at test time, which stems from the inclusion of unseen compositions recombined from attributes and objects. To overcome the challenge, we propose a novel approach that accumulates comprehensive knowledge in both textual and visual modalities from unsupervised data to update multimodal prototypes at test time. Building on this, we further design an adaptive update weight to control the degree of prototype adjustment, enabling the model to flexibly adapt to distribution shift during testing. Moreover, a dynamic priority queue is introduced that stores high-confidence images to acquire visual prototypes from historical images for inference. Since the model tends to favor compositions already stored in the queue during testing, we warm-start the queue by initializing it with training images for visual prototypes of seen compositions and generating unseen visual prototypes using the mapping learned between seen and unseen textual prototypes. Considering the semantic consistency of multimodal knowledge, we align textual and visual prototypes by multimodal collaborative representation learning. To provide a more reliable evaluation for CZSL, we introduce a new benchmark dataset, C-Fashion, and refine the widely used but noisy MIT-States dataset. Extensive experiments indicate that our approach achieves state-of-the-art performance on four benchmark datasets under both closed-world and open-world settings. The source code and datasets are available at https://github.com/xud-yan/WARM-CAT .

零样本学习多模态测试时学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。