arXiv:2410.02678cs.CLcs.AI2024-10被引 43

用文本模型自监督训练语音大模型,无需指令数据和标注回复。

Distilling an End-to-End Voice Assistant Without Instruction Training Data

  • 用文本大模型对语音转写结果生成回复,作为自监督信号训练语音大模型。
  • 在口语问答、分类、翻译任务上表现优秀,用户偏好测试胜率达72%。
  • 仅需顶尖模型1/100的算力,适合资源有限但追求高效率的团队。

语音助手如Siri和Google Assistant通常分别建模音频与文本,导致语音信息丢失且系统复杂。近期基于监督微调(SFT)训练的端到端语音大模型(Speech LLMs)虽提升性能,却会遗忘纯文本大模型的能力。本文提出一种新范式:无需指令数据,利用文本大模型对语音转录文本的响应作为自监督信号训练语音大模型。该过程无需人工标注回复。实验表明,所提出的蒸馏语音助手(DiVA)在口语问答、分类和翻译任务上具备良好泛化能力。此外,用户偏好测试中,DiVA以72%的胜率超越当前领先模型Qwen 2 Audio,且训练计算量不足其1/100。

原文摘要 · Abstract (English)

Voice assistants, such as Siri and Google Assistant, typically model audio and text separately, resulting in lost speech information and increased complexity. Recent efforts to address this with end-to-end Speech Large Language Models (LLMs) trained with supervised finetuning (SFT) have led to models ``forgetting" capabilities from text-only LLMs. Our work proposes an alternative paradigm for training Speech LLMs without instruction data, using the response of a text-only LLM to transcripts as self-supervision. Importantly, this process can be performed without annotated responses. We show that our Distilled Voice Assistant (DiVA) generalizes to Spoken Question Answering, Classification, and Translation. Furthermore, we show that DiVA better meets user preferences, achieving a 72\% win rate compared with state-of-the-art models like Qwen 2 Audio, despite using $>$100x less training compute.

语音大模型自监督高效训练端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。