用AI反馈强化学习提升语音模型语义理解能力
Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
- 通过生成多段语音并基于语义度量构建偏好数据
- 在多个语音基准上达到当前最优性能
- 适合关注语音生成语义质量的研究者
尽管无文本语音语言模型(SLMs)在端到端语音到语音建模中展现出潜力,但在语义连贯性和相关性方面仍落后于基于文本的大语言模型(LLMs)。本文提出Align-SLM框架,借鉴人工智能反馈强化学习(RLAIF)的偏好优化思想,增强SLMs的语义理解能力。方法从给定提示生成多个语音延续,并利用语义指标构建直接偏好优化(DPO)所需偏好数据。我们在ZeroSpeech 2021基准下评估词法与句法建模能力,使用语音版StoryCloze数据集评估语义连贯性,并结合GPT4-o评分与人工评估等指标进行测试。实验表明,该方法在多数基准上实现当前SLMs最佳表现,凸显了偏好优化对提升语音模型语义质量的重要性。
原文摘要 · Abstract (English)
While textless Spoken Language Models (SLMs) have shown potential in end-to-end speech-to-speech modeling, they still lag behind text-based Large Language Models (LLMs) in terms of semantic coherence and relevance. This work introduces the Align-SLM framework, which leverages preference optimization inspired by Reinforcement Learning with AI Feedback (RLAIF) to enhance the semantic understanding of SLMs. Our approach generates multiple speech continuations from a given prompt and uses semantic metrics to create preference data for Direct Preference Optimization (DPO). We evaluate the framework using ZeroSpeech 2021 benchmarks for lexical and syntactic modeling, the spoken version of the StoryCloze dataset for semantic coherence, and other speech generation metrics, including the GPT4-o score and human evaluation. Experimental results show that our method achieves state-of-the-art performance for SLMs on most benchmarks, highlighting the importance of preference optimization to improve the semantics of SLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。