用连续对话+双阶段训练,让语音聊天机器人反应更快更自然。
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- 用连续语句和停顿模拟人类对话节奏
- 训练时交替让话语领先或落后音频,实现语义对齐
- 70亿参数模型,少数据训练仍效果好
全双工对话模型旨在同时听和说,实现对动态用户输入的快速响应。现有方法在每个时间步合并多个通道,但通常将文本话语拆分为词级以对齐音频流,损害语言建模能力。为此,我们提出“连续话语”——由连续语句和“等待”间隔组成,模仿人类对话中的认知行为。研究发现,恰当的训练范式对语义对齐至关重要。为此,我们设计了“双阶段训练”策略:在不同训练阶段交替改变话语相对于音频的位置(领先或滞后)。结合连续话语与双阶段训练策略,我们构建了70亿参数的全双工语音聊天机器人FLM-Audio。实验表明,该模型在响应质量与对话体验上表现优异,且显著降低训练数据需求。
原文摘要 · Abstract (English)
Full-duplex dialog models aim to listen and speak simultaneously, delivering rapid responses to dynamic user input. Among different solutions to full-duplexity, a native solution merges multiple channels in each time step, achieving the lowest latency. However, prevailing designs break down the textual monologue sentences for word-level alignment with audio streams, which degrades language modeling abilities. To help address this issue, we introduce "contiguous monologues", which are composed by continuous sentences and "waiting" intervals, mimicking human-like cognitive behavior in dialogs. We find a proper training paradigm to be critical for semantically aligning contiguous monologues with audio. To this end, we develop a "dual" training paradigm that alternates the position of the monologues, either leading or trailing the audio, across different training stages. A combination of our contiguous monologue and dual training strategy is applied in developing FLM-Audio, our 7B spoken dialog chatbot with native full-duplexity. As confirmed by experimental results, FLM-Audio achieves superior response qualities and chatting experiences while requiring significantly less training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。