构建可扩展的多轮音频预处理流水线,支持全双工语音模型训练
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
- 设计开源流水线,支持多说话人实时对话数据处理
- 解决重叠对话与回音确认等自然对话难题
- 适合研究全双工语音交互与多说话人建模的团队
随着AI范式从文本大模型转向语音语言模型(SLMs),对支持实时、自然人机交互的全双工系统的需求日益增长。然而,这类模型的发展受限于高质量多说话人对话数据的匮乏,现有大规模资源多为单说话人或数据量有限。自然对话中复杂的动态行为,如语音重叠和回音确认,仍难以有效处理,标准处理流程常因语音分离错误和自动语音识别幻觉而失效。为此,我们提出一种鲁棒且可扩展的开源数据处理流水线,专为全双工语音语言模型设计。
原文摘要 · Abstract (English)
As the paradigm of AI shifts from text-based LLMs to Speech Language Models (SLMs), there is a growing demand for full-duplex systems capable of real-time, natural human-computer interaction. However, the development of such models is constrained by the scarcity of high-quality, multi-speaker conversational data, as existing large-scale resources are predominantly single-speaker or limited in volume. Addressing the complex dynamics of natural dialogue, such as overlapping and back-channeling remains a challenge, with standard processing pipelines suffering from diarization errors and ASR hallucinations. To bridge this gap, we present a robust and scalable open-source data processing pipeline designed for full-duplex model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。