用GRPO优化语音感知大模型,提升开放格式语音理解能力
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
- 基于组相对策略优化,用BLEU做奖励信号训练模型
- 在问答和语音翻译任务上优于传统微调方法
- 支持离策略样本,适合研究生成式语音理解的学者
本文提出一种基于组相对策略优化(GRPO)的方法,用于在开放格式语音理解任务(如口语问答和自动语音翻译)上训练语音感知大语言模型(SALLMs)。SALLMs在语音理解任务中表现优异。GRPO因其在大模型训练中的高效性日益受到关注,已有研究将其应用于SALLMs,但主要集中在多项选择类任务。本文聚焦更具生成能力挑战的开放格式任务。方法采用GRPO结合BLEU作为奖励信号优化SALLMs,实验表明其在多个关键指标上超越标准监督微调(SFT)。此外,探讨了在GRPO中引入离策略样本的可能性,为未来改进与研究提供方向。
原文摘要 · Abstract (English)
In this paper, we introduce a Group Relative Policy Optimization (GRPO)-based method for training Speech-Aware Large Language Models (SALLMs) on open-format speech understanding tasks, such as Spoken Question Answering and Automatic Speech Translation. SALLMs have proven highly effective for speech understanding tasks. GRPO has recently gained traction for its efficiency in training LLMs, and prior work has explored its application to SALLMs, primarily in multiple-choice tasks. Building on this, we focus on open-format tasks that better reflect the generative abilities of the models. Our approach leverages GRPO with BLEU as the reward signal to optimize SALLMs, and we demonstrate empirically that it surpasses standard SFT across several key metrics. Finally, we explore the potential of incorporating off-policy samples within GRPO for these tasks, highlighting avenues for further improvement and further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。