用强化学习对齐语音大模型行为,仅靠自生成数据就达到顶尖表现。
Enhancing Speech Large Language Models through Reinforced Behavior Alignment
- 用强大教师模型自动生成高保真对齐数据,避免人工标注。
- 在指令遵循任务中超越传统蒸馏方法,提升语音大模型能力。
- 可无缝扩展到语音问答与语音翻译,适合语音交互研究者。
近期大型语言模型(LLMs)的发展推动了将其语言能力从文本扩展到其他模态的研究,催生了能够处理语音或文本输入的语音大模型(SpeechLMs)。然而,由于跨模态差异,这些SpeechLMs在指令遵循能力上仍显著落后于文本型大模型,尤其是在面对动态多变的用户语音时。为解决此问题,本文提出一种名为强化行为对齐(RBA)的框架,旨在增强SpeechLM的语言生成能力。RBA不依赖人工标注的监督微调,而是通过一个强大的教师模型采用自合成方法生成大量高质量对齐数据,随后使用基于强化学习的方法将SpeechLM的行为与教师模型对齐。实验表明,该方法有效提升了SpeechLM的指令遵循能力,性能优于传统蒸馏基线。更重要的是,RBA可无缝扩展至语音问答和语音到文本翻译等任务,在开放基准测试中仅使用自生成数据即达到当前最优水平。
原文摘要 · Abstract (English)
The recent advancements of Large Language Models (LLMs) have spurred considerable research interest in extending their linguistic capabilities beyond text to other modalities, which leads to emergence of speech-based LLMs (SpeechLMs) with capability of processing user request in either speech or textual formats. However, owing to inter-modal discrepancies, these SpeechLMs still exhibit a significant performance gap compared to their text-based LLM counterparts in instruction-following, particularly when confronted with the dynamic and variable nature of user speech. To address this challenge, this paper introduces a framework termed Reinforced Behavior Alignment (RBA), designed to bolster the language generation proficiency of SpeechLMs. Instead of relying on supervised fine-tuning from human annotations, RBA employs a self-synthesis methodology to generate extensive, high-fidelity alignment data by a powerful teacher LLM. Then SpeechLMs is aligned its behavior with that of a teacher using a reinforcement learning-based approach. Experimental results demonstrate that this method effectively enhances the instruction-following capabilities of SpeechLMs that outperform conventional distillation baselines. Crucially, we demonstrate that RBA can be seamlessly extended to tasks such including spoken question answering and speech-to-text translation, attaining state-of-the-art performance on open benchmarks with only self-generated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。