arXiv:2506.20986cs.CV2025-06

通过专家协作提升组合零样本学习的语义对齐能力

EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning

  • 用多个专家分别处理不同语义子集,实现细粒度特征学习
  • 在三个基准上闭开世界设置下均超越现有方法
  • 适合研究组合泛化与跨模态对齐的学者

组合零样本学习(CZSL)旨在基于已学基础概念识别未知状态-物体组合。现有方法通常采用简单的组合原型映射提取基础特征,难以应对可划分为不同语义子集的个体。此外,全对一的跨模态基础特征匹配忽略了相同状态或物体间的组合差异,限制了图像与组合之间的精细对齐。本文提出EVA框架,即基于混合专家的语义变体对齐方法。通过引入领域专家自适应机制,利用多个专家实现标记感知学习,构建高质量的基础表征。为进一步实现精确的组合泛化,提出语义变体对齐策略,选择语义相关的表示进行图像-基础特征匹配。该方法在三个主流基准上,于闭世界和开世界设置下均显著优于现有最先进方法,验证了所提思路的有效性。

原文摘要 · Abstract (English)

Compositional Zero-Shot Learning (CZSL) investigates compositional generalization capacity to recognize unknown state-object pairs based on learned primitive concepts. Existing CZSL methods typically derive primitives features through a simple composition-prototype mapping, which is suboptimal for a set of individuals that can be divided into distinct semantic subsets. Moreover, the all-to-one cross-modal primitives matching neglects compositional divergence within identical states or objects, limiting fine-grained image-composition alignment. In this study, we propose EVA, a Mixture-of-Experts Semantic Variant Alignment framework for CZSL. Specifically, we introduce domain-expert adaption, leveraging multiple experts to achieve token-aware learning and model high-quality primitive representations. To enable accurate compositional generalization, we further present semantic variant alignment to select semantically relevant representation for image-primitives matching. Our method significantly outperforms other state-of-the-art CZSL methods on three popular benchmarks in both closed- and open-world settings, demonstrating the efficacy of the proposed insight.

零样本学习组合泛化专家模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。