arXiv:2505.02625cs.CLcs.AI2025-05ACL被引 90

用少量数据训练出实时语音对话大模型,性能超越大量数据训练的旧模型。

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

  • 基于Qwen2.5架构,融合语音编码器与自回归流式解码器。
  • 仅用20万条语音对话数据,就达到顶尖语音问答表现。
  • 适合需要低资源、高实时性的语音交互系统研发者。

实时、智能且自然的语音交互是下一代人机交互的核心。本文提出LLaMA-Omni 2,一系列参数量从0.5B到14B的语音语言模型(SpeechLM),可实现高质量实时语音交互。该模型基于Qwen2.5系列,集成语音编码器与自回归流式语音解码器。尽管仅在20万条多轮语音对话样本上训练,其在多个语音问答与语音指令跟随基准测试中表现优异,超越此前最先进的语音语言模型GLM-4-Voice(后者训练数据达数百万小时)。

原文摘要 · Abstract (English)

Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs). In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving high-quality real-time speech interaction. LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder. Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-the-art SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data.

语音对话大模型实时交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。