arXiv:2606.02615eess.AScs.AI2026-06被引 1

让语音大模型学会用少量示例快速适应新任务

FSA-GRPO: Teaching Auditory LLMs to Use Few-Shot Demonstrations

论文配图:FSA-GRPO: Teaching Auditory LLMs to Use Few-Shot Demonstrations
图 1 · 摘自论文原文
  • 用强化学习设计奖励机制,引导模型利用少样本示例
  • 仅用2000条成人语音数据,儿童语音识别错误率降53.9%
  • 无需领域内训练,对低资源语言和语音翻译同样有效

少样本提示为将语音大模型适配至低资源任务(如儿童语音识别)提供了有效方式。然而,大多数语音大模型未显式训练以在示范条件格式下进行推理,限制了其在上下文学习(ICL)中的潜力。为此,我们提出少样本感知的GRPO(FSA-GRPO),一种基于强化学习的后训练方法,通过特殊设计的奖励机制,促使模型充分利用少样本示范,从而增强其少样本适应能力。值得注意的是,仅使用2000条高资源成人语音数据进行训练,即可显著提升模型的通用少样本适应能力,在儿童语音识别中实现53.9%相对词错误率降低(无需任何领域内训练),同时在多语言语音识别(包括低资源语言)、语音翻译和音频理解任务中均取得提升。我们进一步研究了数据选择及辅助奖励的权重与相似度阈值,确定了有效的训练方案。实验表明,当无法获取或使用领域内数据时,FSA-GRPO优于直接在相关域外数据上微调。

原文摘要 · Abstract (English)

Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large language models are not explicitly trained to perform inference in this demonstration-conditioned format, limiting the extent to which they can benefit from In-Context Learning (ICL). To address this limitation, we introduce Few-Shot Aware GRPO (FSA-GRPO), an RL-based post-training recipe that uses a specially designed reward to encourage the model to leverage few-shot demonstrations, thereby strengthening its few-shot adaptation ability. Notably, training with only 2k high-resource adult ASR utterances improves the model's general few-shot adaptation ability, yielding gains not only in children's speech recognition (53.9% relative WER reduction without any in-domain training) but also in multilingual ASR (including low-resource languages), speech translation, and audio understanding. We further study data selection and the weight and similarity cutoffs of the auxiliary reward to identify an effective training recipe. Our experiments show that when in-domain data are unavailable or cannot be used for training, FSA-GRPO is more effective than direct tuning on related out-of-domain data.

语音大模型少样本学习强化学习语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。