arXiv:2607.04941cs.CLcs.SD2026-07

构建大规模双工对话语音数据集,支持更自然的语音对话建模。

DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling

  • 从播客自动提取双人对话,分离每位说话人音频流。
  • 生成超40万小时英日语双工对话数据,含真实对话轮换特征。
  • 适合研究对话系统、语音分离与多说话人建模的团队使用。

全双工口语对话模型需以独立声道表示每位说话人,但现有大规模公开语音语料多为单声道,不适用于此类建模。本文提出 DuplexChat 开源语料库及 DuplexChat-Pipe 数据构建管道,可从公开播客流中提取并处理成双说话人分离的全双工对话语音。该管道通过语言过滤、音频获取与清洗、基于说话人归因的对话片段提取,并结合语音分离与修复技术,为每位说话人生成独立音轨。运行后获得覆盖 282,634 小时英语和 132,723 小时日语的分离式口语对话数据。对 DuplexChat 的分析表明,其包含人类对话中典型的对话轮换特征。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora are mostly monaural, making them unsuited for SDLM training. We present DuplexChat, an open-source corpus for full-duplex spoken dialogue models, and DuplexChat-Pipe, a pipeline for constructing speaker-separated full-duplex dialogue speech from public podcast feeds. DuplexChat-Pipe filters language-specific podcast feeds, retrieves and cleans episode audio, extracts diarization-guided two-speaker dialogue clips, and applies speech separation and restoration to produce one channel per speaker. Running this pipeline yields a speaker-separated spoken dialogue corpus covering 282,634 hours of English and 132,723 hours of Japanese. Analysis results on DuplexChat show that it contains turn-taking dynamics present in human dialogues.

语音对话数据构建说话人分离全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。