arXiv:2412.15649eess.AS2024-12ACL被引 81

单阶段训练实现可控制音色的实时语音对话系统

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

  • 用语义标记建模语言,分离说话人信息至声码器,实现零样本音色控制
  • 每步预测分组语义标记,序列长度大幅缩短,训练仅需15小时
  • 无需预训练即可实现多轮多语言对话,适合低资源语音交互场景

近期进展表明端到端实时语音对话系统具有低延迟和高质量的潜力。本文提出SLAM-Omni,一种支持音色可控、单阶段训练的端到端语音交互系统。该系统通过语义标记建模语言,并将说话人信息解耦至声码器,实现零样本音色控制。通过每步预测分组语音语义标记,显著缩短音频标记序列长度,加速训练与推理。此外,提出历史文本提示机制以压缩对话历史,提升多轮交互效率。综合评估显示,SLAM-Omni在相似规模模型中表现更优,仅需在4块GPU上用有限数据训练15小时。其为首个实现竞争性性能的单阶段训练语音对话系统,无需在语音合成(TTS)或语音识别(ASR)任务上进行预训练。进一步实验验证了其在更大数据集上的多语言与多轮对话能力。

原文摘要 · Abstract (English)

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system with single-stage training. SLAM-Omni achieves zero-shot timbre control by modeling spoken language with semantic tokens and decoupling speaker information to a vocoder. By predicting grouped speech semantic tokens at each step, our method significantly reduces the sequence length of audio tokens, accelerating both training and inference. Additionally, we propose historical text prompting to compress dialogue history, facilitating efficient multi-round interactions. Comprehensive evaluations reveal that SLAM-Omni outperforms prior models of similar scale, requiring only 15 hours of training on 4 GPUs with limited data. Notably, it is the first spoken dialogue system to achieve competitive performance with a single-stage training approach, eliminating the need for pre-training on TTS or ASR tasks. Further experiments validate its multilingual and multi-turn dialogue capabilities on larger datasets.

语音交互音色控制端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。