arXiv:2603.06281cs.CV2026-03

解决生成式零样本学习中属性与视觉不匹配的问题

Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning

  • 建模类别属性分布,生成更贴合实例的属性
  • 提升语义与视觉特征对齐,准确率提高4.7%~6.1%
  • 可插件式增强现有生成式零样本方法

生成式零样本学习(ZSL)通过语义条件将已见类知识迁移至未见类,但面临两大挑战:(1)类别级属性难以捕捉类内实例的视觉差异,导致类-实例差距;(2)语义与视觉特征分布存在显著错配,表现为类间相关性,引发语义-视觉域差距。为此,提出属性分布建模与语义-视觉对齐(ADiVA)方法,联合建模属性分布并显式对齐语义与视觉特征。ADiVA包含两个模块:属性分布建模(ADM)模块为每类学习可迁移的属性分布,并为未见类采样实例级属性;视觉引导对齐(VGA)模块优化语义表示以更好反映视觉结构。在三个主流基准数据集上的实验表明,ADiVA显著优于当前最优方法(如在AWA2和SUN上分别提升4.7%和6.1%)。此外,该方法可作为插件提升现有生成式零样本学习模型性能。

原文摘要 · Abstract (English)

Generative zero-shot learning (ZSL) synthesizes features for unseen classes, leveraging semantic conditions to transfer knowledge from seen classes. However, it also introduces two intrinsic challenges: (1) class-level attributes fails to capture instance-specific visual appearances due to substantial intra-class variability, thus causing the class-instance gap; (2) the substantial mismatch between semantic and visual feature distributions, manifested in inter-class correlations, gives rise to the semantic-visual domain gap. To address these challenges, we propose an Attribute Distribution Modeling and Semantic-Visual Alignment (ADiVA) approach, jointly modeling attribute distributions and performing explicit semantic-visual alignment. Specifically, our ADiVA consists of two modules: an Attribute Distribution Modeling (ADM) module that learns a transferable attribute distribution for each class and samples instance-level attributes for unseen classes, and a Visual-Guided Alignment (VGA) module that refines semantic representations to better reflect visual structures. Experiments on three widely used benchmark datasets demonstrate that ADiVA significantly outperforms state-of-the-art methods (e.g., achieving gains of 4.7% and 6.1% on AWA2 and SUN, respectively). Moreover, our approach can serve as a plugin to enhance existing generative ZSL methods.

零样本学习属性建模语义对齐生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。