arXiv:2606.15751cs.SDcs.LG2026-06中稿 · INTERSPEECH 2026

在音频编码器中加入可训练提示,提升少样本音频分类性能

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

论文配图:Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models
图 1 · 摘自论文原文
  • 在音频编码器中引入可学习提示,捕捉任务特异性声学特征
  • 在11个数据集上验证,结合文本提示后少样本效果显著提升
  • 无需修改模型结构,可作为插件模块直接集成使用

音频语言模型(ALMs)在零样本音频分类任务中表现优异,通过对齐音频波形与文本实现。近期研究聚焦于优化文本提示以提升下游性能,但多数方法仅关注文本编码器,忽略了音频编码器中可学习提示的潜力。本文提出一种新框架,在音频编码器中引入可训练提示,以捕捉任务相关的声学特征。实验表明,将音频侧提示学习与现有文本侧方法结合,能有效提升少样本适应能力。在11个数据集上的广泛实验显示,该方法作为即插即用模块与现有文本提示调优结合时,通常带来性能提升。结果表明,显式调节音频表示空间可有效补充仅依赖文本提示的方法。代码已开源:https://github.com/hyebin-c/aspl。

原文摘要 · Abstract (English)

Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance focus on learning optimal text prompts. However, previous approaches focus on the text encoder, leaving the potential of learnable prompts within the audio encoder unexplored. In this paper, we propose a novel framework that introduces trainable prompts into the audio encoder to capture task-specific acoustic features. We demonstrate that integrating audio-side prompt learning with existing text-side approaches enhances few-shot adaptation. Through extensive experiments across 11 datasets show that integrating our method as a plug-and-play module alongside existing text prompt tuning generally leads to performance improvements. These findings suggest that explicitly modulating the audio representation space effectively complements text-only prompting approaches. The code is available at https://github.com/hyebin-c/aspl.

音频模型少样本学习提示工程自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。