arXiv:2411.18138eess.AScs.AI2024-11被引 42

无需编码器的全双工语音模型,可边说边听,支持自然对话。

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

  • 用嵌入向量替代编码器,实现端到端语音理解与生成。
  • 在语音识别、降噪、问答等任务上表现优异,支持实时换言与回声消除。
  • 适合开发真正自然的语音助手和全双工对话系统。

全双工多模态大语言模型为语音理解和生成任务提供了统一框架,使人机对话更自然流畅。与传统分模块系统不同,多模态大模型作为单一端到端模型运行,避免了组件间误差传播,并充分利用输入语音信号中的丰富非语言信息。我们提出 SALMONN-omni,一种无需编码器的全双工语音理解与生成模型,可在说话时同时监听自身生成语音和背景声音。为此,我们设计了一种新型双工对话框架,引入“思考”机制,依赖嵌入向量而非编码器(量化语音与音频标记)实现异步文本与语音生成。实验表明,SALMONN-omni 在多种流式语音任务中表现出色,包括语音识别、语音增强和口语问答。此外,该模型在换言、打断和回声消除场景中表现卓越,展现出构建鲁棒全双工对话系统的潜力。据我们所知,SALMONN-omni 是首个此类无编码器模型。完整技术报告及模型检查点即将发布。

原文摘要 · Abstract (English)

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike traditional modularised conversational AI systems, which separate speech recognition, understanding, and text-to-speech generation into distinct components, multimodal LLMs operate as single end-to-end models. This streamlined design eliminates error propagation across components and fully leverages the rich non-verbal information embedded in input speech signals. We introduce SALMONN-omni, a codec-free, full-duplex speech understanding and generation model capable of simultaneously listening to its own generated speech and background sounds while speaking. To support this capability, we propose a novel duplex spoken dialogue framework incorporating a ``thinking'' mechanism that facilitates asynchronous text and speech generation relying on embeddings instead of codecs (quantized speech and audio tokens). Experimental results demonstrate SALMONN-omni's versatility across a broad range of streaming speech tasks, including speech recognition, speech enhancement, and spoken question answering. Additionally, SALMONN-omni excels at managing turn-taking, barge-in, and echo cancellation scenarios, establishing its potential as a robust prototype for full-duplex conversational AI systems. To the best of our knowledge, SALMONN-omni is the first codec-free model of its kind. A full technical report along with model checkpoints will be released soon.

全双工语音生成大模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。