arXiv:2411.00774cs.SDcs.AI2024-11ICML被引 176

冻结大模型实现低延迟语音对话,保持智能水平不下降。

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

  • 语音输入输出通过三阶段训练接入冻结文本大模型。
  • 仅用6万条对话数据,在8块GPU上实现端到端语音响应。
  • 支持双向对话,适合资源有限但需自然语音交互的场景。

快速发展的大语言模型(LLMs)带来了巨大的智能应用潜力。特别是GPT-4o出色的双工语音交互能力,为用户带来了惊艳体验。近期已有研究提出多种多模态LLM,实现用户与代理之间的语音对话语音对话。本文提出一种新型语音-文本多模态LLM架构——Freeze-Omni。核心贡献在于,语音输入与输出模态可轻松连接至文本大模型,并在整个训练过程中保持大模型参数冻结。我们设计了三阶段训练策略,仅使用60,000条多轮文本问答数据及文本-语音配对数据(如ASR和TTS数据),在8张GPU上即可使Freeze-Omni具备语音对话语音对话能力。同时,有效保证了其语音模态的智能水平与文本大模型基座相当,且实现低延迟端到端语音响应。此外,我们还通过多任务学习方法实现了双工对话能力,使系统表现出更自然的对话风格。综上,Freeze-Omni在冻结大模型条件下,具备基于多模态LLM进行语音对话语音对话的巨大潜力,避免了因数据与训练资源有限导致的灾难性遗忘问题。

原文摘要 · Abstract (English)

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users. Researchers have recently proposed several multi-modal LLMs in this direction that can achieve user-agent speech-to-speech conversations. This paper proposes a novel speech-text multimodal LLM architecture called Freeze-Omni. Our main contribution is that the speech input and output modalities can be easily connected to a textual LLM while keeping the LLM's parameters frozen throughout the training process. We design a three-stage training strategy for modeling both the speech input and output, enabling Freeze-Omni to obtain speech-to-speech conversation ability using text-speech paired data (such as ASR and TTS data) and only 60,000 multi-round text Q&A data on 8 GPUs. Moreover, we can effectively ensure that the intelligence of the Freeze-Omni in the speech modality is at the same level compared with that in the text modality of its backbone LLM, while achieving low latency end-to-end spoken response. In addition, we also designed a method to achieve duplex dialogue ability through multi-task training, giving Freeze-Omni a more natural style of dialogue ability between users and agents. In summary, Freeze-Omni holds great potential to conduct speech-to-speech dialogue based on a multimodal LLM under the condition of a frozen LLM, avoiding the catastrophic forgetting problem caused by limited data and training resources.

语音对话多模态冻结模型低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。