提出简单方法区分音频视觉数据中的已见与未见样本,提升零样本学习性能。
Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning
- 利用类别专属逻辑值和类无关特征子空间实现无需额外训练的分布外检测
- 在三个数据集上达到新的最优效果,显著提升零样本与广义零样本学习性能
- 适合关注跨模态零样本学习、数据分布偏移问题的研究者
零样本学习(ZSL)通过挖掘辅助类别信息实现从已见类别到未见类别的知识迁移,是极具前景但极具挑战的研究方向。在该领域中,音视频广义零样本学习(AV-GZSL)因音频、视频与自然语言三模态间复杂关系而尤为困难却极具研究价值。然而,现有基于嵌入和生成的AV-GZSL方法普遍面临领域偏移问题。为此,本文提出一种极简的分布外(OOD)检测方法——EZ-AVOOD,通过在初始阶段区分已见与未见样本以缓解偏差。该方法利用类别专属逻辑值和类无关特征子空间中的内在判别信息,无需训练额外的OOD检测网络即可实现有效的已见-未见分离。随后采用两个专家模型分别对已见和未见样本进行分类。相比现有最先进方法,本模型在三个音视频数据集上均取得更优的ZSL与GZSL性能,成为新SOTA,充分证明了所提方法的有效性。
原文摘要 · Abstract (English)
Zero-shot Learning(ZSL) attains knowledge transfer from seen classes to unseen classes by exploring auxiliary category information, which is a promising yet difficult research topic. In this field, Audio-Visual Generalized Zero-Shot Learning~(AV-GZSL) has aroused researchers' great interest in which intricate relations within triple modalities~(audio, video, and natural language) render this task quite challenging but highly research-worthy. However, both existing embedding-based and generative-based AV-GZSL methods tend to suffer from domain shift problem a lot and we propose an extremely simple Out-of-distribution~(OOD) detection based AV-GZSL method~(EZ-AVOOD) to further mitigate bias problem by differentiating seen and unseen samples at the initial beginning. EZ-AVOOD accomplishes effective seen-unseen separation by exploiting the intrinsic discriminative information held in class-specific logits and class-agnostic feature subspace without training an extra OOD detector network. Followed by seen-unseen binary classification, we employ two expert models to classify seen samples and unseen samples separately. Compared to existing state-of-the-art methods, our model achieves superior ZSL and GZSL performances on three audio-visual datasets and becomes the new SOTA, which comprehensively demonstrates the effectiveness of the proposed EZ-AVOOD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。