提出可插拔的实时对话状态预测模块,提升全双工语音交互效率。
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
- 利用语音转写文本信息实现语义级语音活动检测
- 低延迟下对话轮次管理与响应速度优于现有模型
- 适合需要实时交互的智能语音系统开发者
近期语音对话系统进展推动了类人全双工语音交互的关注。然而,我们对这一领域的综述揭示了训练数据获取难、灾难性遗忘和可扩展性有限等挑战。本文提出SoulX-Duplug,一种面向全双工语音对话系统的即插即用流式状态预测模块。通过联合进行流式语音识别(ASR),SoulX-Duplug显式利用文本信息识别用户意图,有效充当语义级语音活动检测(VAD)。为促进公平评估,我们引入SoulX-Duplug-Eval,扩展了广泛使用的基准测试集以提升双语覆盖。实验结果表明,SoulX-Duplug实现了低延迟的流式对话状态控制,基于该模块构建的系统在整体轮次管理与延迟性能上优于现有全双工模型。SoulX-Duplug与SoulX-Duplug-Eval已开源。
原文摘要 · Abstract (English)
Recent advances in spoken dialogue systems have brought increased attention to human-like full-duplex voice interactions. However, our comprehensive review of this field reveals several challenges, including the difficulty in obtaining training data, catastrophic forgetting, and limited scalability. In this work, we propose SoulX-Duplug, a plug-and-play streaming state prediction module for full-duplex spoken dialogue systems. By jointly performing streaming ASR, SoulX-Duplug explicitly leverages textual information to identify user intent, effectively serving as a semantic VAD. To promote fair evaluation, we introduce SoulX-Duplug-Eval, extending widely used benchmarks with improved bilingual coverage. Experimental results show that SoulX-Duplug enables low-latency streaming dialogue state control, and the system built upon it outperforms existing full-duplex models in overall turn management and latency performance. We have open-sourced SoulX-Duplug and SoulX-Duplug-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。