arXiv:2409.09213eess.AScs.CL2024-09被引 20

用声音特征描述提升零样本音频分类效果

ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds

  • 用声音的本征特征重写提示,增强模型对真实场景声音的理解
  • 在多个数据集上零样本分类性能提升1%-18%,最高超基线55%
  • 适合需要高精度零样本音频分类的研究者与应用开发者

开放词汇音频-语言模型(如CLAP)通过自然语言提示实现任意类别零样本音频分类(ZSAC),具有广阔前景。本文提出一种简单但高效的方法改进基于CLAP的ZSAC。不同于传统使用抽象类别标签(如“管风琴的声音”)的提示方式,我们采用包含声音固有描述性特征的多样化上下文提示(如“管风琴深沉回响的音色充满大教堂”)。为此,我们首先提出ReCLAP:一个使用重写音频描述训练的CLAP模型,其描述每个声音事件的独特判别特征。ReCLAP在多模态音频-文本检索和零样本音频分类任务上均优于所有基线。进一步地,为提升基于ReCLAP的零样本分类性能,我们提出提示增强策略:不再使用手工模板提示,而是为数据集中每个唯一标签生成定制化提示,先描述声音事件,再将其置于多样场景中。该方法使ReCLAP在ZSAC上性能提升1%-18%,且全面超越所有基线1%-55%。

原文摘要 · Abstract (English)

Open-vocabulary audio-language models, like CLAP, offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper, we propose a simple but effective method to improve ZSAC with CLAP. Specifically, we shift from the conventional method of using prompts with abstract category labels (e.g., Sound of an organ) to prompts that describe sounds using their inherent descriptive features in a diverse context (e.g.,The organ's deep and resonant tones filled the cathedral.). To achieve this, we first propose ReCLAP, a CLAP model trained with rewritten audio captions for improved understanding of sounds in the wild. These rewritten captions describe each sound event in the original caption using their unique discriminative characteristics. ReCLAP outperforms all baselines on both multi-modal audio-text retrieval and ZSAC. Next, to improve zero-shot audio classification with ReCLAP, we propose prompt augmentation. In contrast to the traditional method of employing hand-written template prompts, we generate custom prompts for each unique label in the dataset. These custom prompts first describe the sound event in the label and then employ them in diverse scenes. Our proposed method improves ReCLAP's performance on ZSAC by 1%-18% and outperforms all baselines by 1% - 55%.

音频分类零样本学习提示工程CLAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。