arXiv:2412.04083cs.CV2024-12被引 1

提升图像与文本交互,实现更精准的未知组合识别

Unified Framework for Open-World Compositional Zero-shot Learning

  • 通过增强图文模态交互,改进组合零样本学习
  • 在三个数据集上达到最新最好性能,两个超越大视觉语言模型
  • 适合关注跨模态理解与开放世界识别的研究者

开放世界组合零样本学习(OW-CZSL)旨在识别已知元素的新组合。尽管已有方法利用语言知识进行识别,但其图文模态间交互仍较弱。本文提出的方法着重增强图像与文本数据之间的深层交互。此外,引入新模块以缓解推理阶段对所有可能组合进行穷举搜索带来的计算负担。不同于以往仅联合或独立学习组合的方式,本文提出一种混合学习机制,结合两者优势生成最终预测。所提模型在三个数据集上达到当前最优表现,且在两个数据集上超越大型视觉语言模型(LLVM)。

原文摘要 · Abstract (English)

Open-World Compositional Zero-Shot Learning (OW-CZSL) addresses the challenge of recognizing novel compositions of known primitives and entities. Even though prior works utilize language knowledge for recognition, such approaches exhibit limited interactions between language-image modalities. Our approach primarily focuses on enhancing the inter-modality interactions through fostering richer interactions between image and textual data. Additionally, we introduce a novel module aimed at alleviating the computational burden associated with exhaustive exploration of all possible compositions during the inference stage. While previous methods exclusively learn compositions jointly or independently, we introduce an advanced hybrid procedure that leverages both learning mechanisms to generate final predictions. Our proposed model, achieves state-of-the-art in OW-CZSL in three datasets, while surpassing Large Vision Language Models (LLVM) in two datasets.

零样本学习图文交互开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。