arXiv:2409.20007eess.AScs.CL2024-09中稿 · ICASSP 2025被引 56

无需语音指令数据,让语音模型学会理解复杂指令。

DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

论文配图:DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
图 1 · 摘自论文原文
  • 自动生成语音文本对数据,注入语音语调理解能力。
  • 在Dynamic-SUPERB和AIR-Bench-Chat上表现优异,无需语音微调。
  • 可执行链式推理与格式化输出,适合需要语音指令的场景。

近期端到端语音语言模型(SLMs)通过融合预训练语音模型,拓展了大语言模型(LLMs)的能力。然而,这些SLMs通常需大量语音指令微调以弥合语音与文本模态的差距,这不仅耗费人力标注,还可能导致原始语言能力的灾难性遗忘。本文提出一种简单高效的自动数据生成流程,通过精心构造语音-文本对,在保留文本LLM原有语言能力的同时,注入语音语调理解能力。实验表明,该模型在无需语音指令微调数据的前提下,即可在Dynamic-SUPERB和AIR-Bench-Chat基准上取得出色性能,具备遵循复杂指令的能力,如特定输出格式和链式思维推理。该方法显著提升了SLMs的通用性与效率,减少对大规模标注数据的依赖,为更高效、更强的语音理解系统开辟新路径。

原文摘要 · Abstract (English)

Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires significant annotation efforts and risks catastrophic forgetting of the original language capabilities. In this work, we present a simple yet effective automatic process for creating speech-text pair data that carefully injects speech paralinguistic understanding abilities into SLMs while preserving the inherent language capabilities of the text-based LLM. Our model demonstrates general capabilities for speech-related tasks without the need for speech instruction-tuning data, achieving impressive performance on Dynamic-SUPERB and AIR-Bench-Chat benchmarks. Furthermore, our model exhibits the ability to follow complex instructions derived from LLMs, such as specific output formatting and chain-of-thought reasoning. Our approach not only enhances the versatility and effectiveness of SLMs but also reduces reliance on extensive annotated datasets, paving the way for more efficient and capable speech understanding systems.

语音语言模型指令跟随零样本语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。