用语音文本交替训练,让语音大模型更泛化
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
- 通过语音与文字配对数据模仿行为,无需大量标注
- 在少样本条件下实现比现有模型更强的任务泛化能力
- 适合想提升语音大模型通用性的研究者和开发者
大语言模型(LLMs)在跨任务上表现出卓越的泛化能力,推动了语音与大模型结合的研究。语音大语言模型(SLLMs)通常采用监督微调来对齐语音与文本模型,但因缺乏广泛任务下的标注语音数据,导致对齐效率低、泛化能力差。为此,我们提出一种仅依赖语音与文本配对数据的多任务‘行为模仿’方法,称为MTBI,通过确保解码器对语音和对应文本生成等效响应,实现更优泛化的SLLM。引入语音-文本交替机制进一步提升对齐效率。我们设计了一个简单基准,用于评估不同模型在提示和任务泛化上的表现。实验表明,所提MTBI在提示与任务泛化上均优于当前最优的SLLMs,且所需监督语音数据更少。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown remarkable generalization across tasks, leading to increased interest in integrating speech with LLMs. These speech LLMs (SLLMs) typically use supervised fine-tuning to align speech with text-based LLMs. However, the lack of annotated speech data across a wide range of tasks hinders alignment efficiency, resulting in poor generalization. To address these issues, we propose a novel multi-task 'behavior imitation' method with speech-text interleaving, called MTBI, which relies solely on paired speech and transcripts. By ensuring the LLM decoder generates equivalent responses to paired speech and text, we achieve a more generalized SLLM. Interleaving is used to further enhance alignment efficiency. We introduce a simple benchmark to evaluate prompt and task generalization across different models. Experimental results demonstrate that our MTBI outperforms SOTA SLLMs on both prompt and task generalization, while requiring less supervised speech data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。