arXiv:2409.00408cs.SDcs.LG2024-09中稿 · International Work…被引 3

用时间注意力提升多标签音频零样本分类准确率

Multi-label Zero-Shot Audio Classification with Temporal Attention

  • 引入时间注意力机制,动态加权音频片段重要性
  • 在AudioSet子集上显著优于均匀聚合特征的方法
  • 适合需要泛化到未见音频类别的研究者

零样本学习模型通过利用辅助信息从已见类别迁移知识,实现对新类别的分类。尽管现有方法多集中于单标签分类任务,本文提出一种多标签零样本音频分类方法。为应对多标签声音分类并泛化至未见类别的挑战,我们引入时间注意力机制,根据声学与语义兼容性为不同音频段分配重要性权重,使模型能聚焦于每类声音最相关的片段,从而捕捉音频样本中各类别主导性的变化。相比不加权的时序聚合特征方法(均等处理所有片段),该方法显著提升多标签零样本分类性能。我们在AudioSet的一个子集上进行了评估,对比了使用均匀聚合特征的零样本模型、零规则基线及本文提出的方法。结果表明,时间注意力有效提升了零样本音频分类在多标签场景下的表现。

原文摘要 · Abstract (English)

Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot learning methods focused on single-label classification tasks, the present study introduces a method to perform multi-label zero-shot audio classification. To address the challenge of classifying multi-label sounds while generalizing to unseen classes, we adapt temporal attention. The temporal attention mechanism assigns importance weights to different audio segments based on their acoustic and semantic compatibility, thus enabling the model to capture the varying dominance of different sound classes within an audio sample by focusing on the segments most relevant for each class. This leads to more accurate multi-label zero-shot classification than methods employing temporally aggregated acoustic features without weighting, which treat all audio segments equally. We evaluate our approach on a subset of AudioSet against a zero-shot model using uniformly aggregated acoustic features, a zero-rule baseline, and the proposed method in the supervised scenario. Our results show that temporal attention enhances the zero-shot audio classification performance in multi-label scenario.

音频分类零样本学习时间注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。