arXiv:2501.00398cs.SDcs.AI2025-01中稿 · SALMA Workshop ICA…被引 3

通过定制化提示词提升音频模型零样本分类性能

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification

  • 基于任务特征生成富含上下文的提示词,替代通用模板
  • 在12个数据集上实现1.23%-16.36%的准确率提升
  • 适合需要快速适配新音频任务的研究者使用

音频-语言模型(ALMs)在零样本音频分类中表现优异,即通过自然语言提示对测试时未见的音频片段进行分类。本文提出TSPE(任务特定提示集成),一种无需训练的硬提示方法,通过为不同音频分类任务定制提示词,显著提升ALMs的零样本性能。不同于通用模板如“汽车的声音”,我们利用标签信息识别合适的音效属性(如“响亮”、“微弱”)和声源(如“隧道”、“街道”),并将这些信息融入提示中。此外,通过在生成的任务特定提示间进行提示集成,进一步增强音文对齐效果。在12个多样化音频分类数据集上的评估表明,相较于原始零样本评测,TSPE使各类ALMs的性能绝对提升1.23%至16.36%。

原文摘要 · Abstract (English)

Audio-language models (ALMs) excel in zero-shot audio classification, a task where models classify previously unseen audio clips at test time by leveraging descriptive natural language prompts. We introduce TSPE (Task-Specific Prompt Ensemble), a simple, training-free hard prompting method that boosts ALEs' zero-shot performance by customizing prompts for diverse audio classification tasks. Rather than using generic template-based prompts like "Sound of a car" we generate context-rich prompts, such as "Sound of a car coming from a tunnel". Specifically, we leverage label information to identify suitable sound attributes, such as "loud" and "feeble", and appropriate sound sources, such as "tunnel" and "street" and incorporate this information into the prompts used by Audio-Language Models (ALMs) for audio classification. Further, to enhance audio-text alignment, we perform prompt ensemble across TSPE-generated task-specific prompts. When evaluated on 12 diverse audio classification datasets, TSPE improves performance across ALMs by showing an absolute improvement of 1.23-16.36% over vanilla zero-shot evaluation.

零样本分类提示工程音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。