用广播业经验与语音偏好数据提升大模型语音适配性
Speechworthy Instruction-tuned Language Models
- 基于广播行业规范设计语音友好提示策略
- 构建20K条语音对比数据集,实现语音适配偏好学习
- 提示+偏好学习可叠加提升,76.2%对话胜率优于基线
当前指令调优语言模型仅使用文本偏好数据训练,难以满足语音等多模态需求。为更好对齐语音领域,本文探索(i)基于广播行业最佳实践的提示策略,及(ii)利用20,000样本的新型语音偏好数据进行偏好学习,该数据通过多种提示生成不同语音适配度的回复,并由听觉标注员对成对回复进行评分。人工与自动评估均表明,提示策略与偏好学习均可提升主流指令调优大模型的语音适配性。有趣的是,二者具有叠加效应:联合使用时在一对一比较中平均获得76.2%的胜率或平局,显著优于基线模型。最后,通过词汇、句法和定性分析揭示了两种方法如何分别提升生成回复的语音适配性。
原文摘要 · Abstract (English)
Current instruction-tuned language models are exclusively trained with textual preference data and thus are often not aligned with the unique requirements of other modalities, such as speech. To better align language models with the speech domain, we explore (i) prompting strategies grounded in radio-industry best practices and (ii) preference learning using a novel speech-based preference data of 20K samples, generated with a wide spectrum of prompts that induce varying dimensions of speech-suitability and labeled by annotators who listen to response pairs. Both human and automatic evaluation show that both prompting and preference learning increase the speech-suitability of popular instruction-tuned LLMs. Interestingly, we find that prompting and preference learning can be additive; combining them achieves the best win rates in head-to-head comparison, resulting in responses that are preferred or tied to the base model in 76.2% of comparisons on average. Lastly, we share lexical, syntactical, and qualitative analyses to showcase how each method contributes to improving the speech-suitability of generated responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。