用声音特征描述提升零样本音频分类效果,无需训练。
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
- 用语音特征描述替代简单标签,提升分类精度
- 在多个主流数据集上达到当前最优零样本表现
- 无需额外训练,适合快速部署的音频识别场景
通过对比学习训练的音文模型可通过自然语言提示实现零样本音频分类,例如使用‘这是一个…的声音’这类提示。本文探索了零样本音频分类中不同的提示模板,发现提示格式对性能影响显著:仅用规范化的类别标签提示即可达到与优化模板甚至提示集成相当的效果。此外,我们引入以声学特征为核心的类别描述,借助大语言模型生成文本描述,有效区分相似类别的声音事件,无需复杂提示工程。实验表明,使用类别描述进行提示可在主要环境声音数据集上实现零样本音频分类的最新成果。值得注意的是,该方法无需额外训练,始终保持完全零样本特性。
原文摘要 · Abstract (English)
Audio-text models trained via contrastive learning offer a practical approach to perform audio classification through natural language prompts, such as "this is a sound of" followed by category names. In this work, we explore alternative prompt templates for zero-shot audio classification, demonstrating the existence of higher-performing options. First, we find that the formatting of the prompts significantly affects performance so that simply prompting the models with properly formatted class labels performs competitively with optimized prompt templates and even prompt ensembling. Moreover, we look into complementing class labels by audio-centric descriptions. By leveraging large language models, we generate textual descriptions that prioritize acoustic features of sound events to disambiguate between classes, without extensive prompt engineering. We show that prompting with class descriptions leads to state-of-the-art results in zero-shot audio classification across major ambient sound datasets. Remarkably, this method requires no additional training and remains fully zero-shot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。