解决音视频零样本学习中模态差异问题,提升未见类识别能力
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
- 设计双注意力机制,分别处理模态间质量与内容差异
- 在多个基准数据集上达到当前最优性能,显著提升未见类别准确率
- 适合关注跨模态对齐与零样本学习的研究者
音视频零样本学习(ZSL)因其能够识别未见类别,在视频分类任务中备受关注。然而,模态不平衡导致模型过度依赖表现较好的模态,削弱了对未见类别的判别能力。现有方法虽尝试通过调整参数梯度缓解该问题,但仍面临两大挑战:(a) 质量差异——不同模态对同一概念提供的信息数量与质量不一;(b) 内容差异——同一模态内样本贡献度差异显著。为此,本文提出一种差异感知注意力网络(DAAN),引入质量差异缓解注意力(QDMA)单元以减少高质量模态中的冗余信息,并设计对比样本级梯度调制(CSGM)模块,通过优化过程与收敛速率量化模态贡献,动态调节梯度幅度,平衡内容差异。实验表明,DAAN在多个基准数据集上达到先进水平,消融实验证明各模块有效性。
原文摘要 · Abstract (English)
Audio-visual Zero-Shot Learning (ZSL) has attracted significant attention for its ability to identify unseen classes and perform well in video classification tasks. However, modal imbalance in (G)ZSL leads to over-reliance on the optimal modality, reducing discriminative capabilities for unseen classes. Some studies have attempted to address this issue by modifying parameter gradients, but two challenges still remain: (a) Quality discrepancies, where modalities offer differing quantities and qualities of information for the same concept. (b) Content discrepancies, where sample contributions within a modality vary significantly. To address these challenges, we propose a Discrepancy-Aware Attention Network (DAAN) for Enhanced Audio-Visual ZSL. Our approach introduces a Quality-Discrepancy Mitigation Attention (QDMA) unit to minimize redundant information in the high-quality modality and a Contrastive Sample-level Gradient Modulation (CSGM) block to adjust gradient magnitudes and balance content discrepancies. We quantify modality contributions by integrating optimization and convergence rate for more precise gradient modulation in CSGM. Experiments demonstrates DAAN achieves state-of-the-art performance on benchmark datasets, with ablation studies validating the effectiveness of individual modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。