通过分层标准化对齐音视频与文本嵌入,提升跨模态零样本分类效果。
On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

- 对音视频与文本嵌入进行Z-score标准化,减少分布差异。
- 在语义、类别和批次三个层次上实现分层对齐,增强嵌入空间结构。
- 在三个基准数据集上表现优异,适合跨模态零样本学习研究者。
音视频广义零样本学习(AV-GZSL)旨在通过融合音频与视觉数据,对已见和未见物体或场景进行分类。现有方法多聚焦于音视频特征的融合或对齐,以生成更具信息量的音视频嵌入;同时,大多数方法仅依赖优化目标来对齐音视频与文本特征,却忽略了音视频与文本模态间的固有分布与结构差异。为此,本文提出一种名为分层标准化嵌入对齐(AHSE)的方法,在共享嵌入空间中实现音视频与文本嵌入的分层对齐。具体而言,首先对融合后的音视频嵌入与文本嵌入进行Z-score标准化,以降低分布不匹配;随后引入分层对齐策略,在语义、类别与批次三个层级最小化差异,构建更鲁棒且结构清晰的嵌入空间。该策略不仅保留了语义与类间关系,还维持了每个批次内的空间一致性。在三个基准数据集:VGGSound-GZSL、UCF-GZSL 和 ActivityNet-GZSL 上的大量实验表明,AHSE 在零样本学习任务中取得了具有竞争力的性能。
原文摘要 · Abstract (English)
Audio-visual Generalized Zero-shot Learning (AV-GZSL) is a challenging task that aims to classify both seen and unseen objects or scenes by integrating data from audio and visual modalities. Recent studies primarily focus on fusing or aligning audio and visual features to generate more informative audio-visual embeddings. Also, aligning the audio-visual and textual features of most existing methods relies solely on the optimization objectives. However, those methods neglect the inherent distributional and structural differences between audio-visual and textual modalities. To address this limitation, we propose a method termed Aligning Hierarchical Standardized Embedding (AHSE), which enables hierarchical alignment of standardized audio-visual and textual embeddings within a shared embedding space. Specifically, we first apply Z-score standardization to the fused audio-visual and textual embeddings to reduce distributional mismatches. We then introduce a hierarchical alignment strategy that minimizes discrepancies at the semantic, class, and batch levels, thereby constructing a more robust and well-structured embedding space. This strategy not only preserves semantic and inter-class relationships but also maintains spatial consistency within each batch. Extensive experiments on three benchmark datasets: VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL, demonstrate that AHSE achieves competitive performance in zero-shot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。