构建真实对话语音数据集,支持双向语音模型训练。
ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion

- 通过分离、重建与扩展三步法,还原真实对话中的交互细节。
- 重建与扩展数据质量高,自然度评分达4.61,对话流畅性接近真人。
- 适合研究双向语音交互、对话系统与语音合成的开发者使用。
全双工语音模型需要保留对话轮次、重叠、打断和回应行为的数据,但这些信号在嘈杂的真实录音中相互纠缠。我们提出 Conversational Voice,一个将真实双人对话片段转化为三种互补训练数据的流水线:(1) 分离阶段恢复各说话人独立音轨、稳定声纹分配、标准转录文本及自然交互时间;(2) 重建阶段从固定源文本生成匹配声音的语音,还原原始发言顺序、停顿与重叠,并添加词级对齐与表达指令;(3) 扩展阶段生成受源上下文、说话人与观察到交互模式约束的新对话。自动声纹验证指标保持强劲,同说话人相似度为0.983–0.991,正向区分度为0.199–0.209。预测语音质量(NISQA MOS)分别为:分离3.56,重建4.41,扩展4.61。基于Gemini的自动评估器对扩展数据的上下文连贯性与对话自然度平均评分分别为4.94/5和4.80/5。扩展与重建在交互特征上高度一致,扩展的轮次、重叠事件、回应行为和打断率分别低4.6%、8.0%、13.2%、16.0%。本工作仅评估数据属性,下游全双工模型训练增益留待后续研究。
原文摘要 · Abstract (English)
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts. (1) Separation recovers speaker-specific tracks with stable speaker assignments, a canonical transcript, and naturally observed interaction timing. (2) Reconstruction generates speech in matched voices from a fixed source transcript, reconstructs the source turn order, pauses, and overlaps, and adds word-level alignment and delivery instructions. (3) Expansion generates new dialogue constrained by the source context, speakers, and observed interaction pattern. Automatic speaker-verification metrics remain strong across stages, with same-speaker similarity of 0.983-0.991 and positive discrimination margins of 0.199-0.209. Predicted speech quality (NISQA MOS) is 3.56 for separation, 4.41 for reconstruction, and 4.61 for expansion. A Gemini-based automatic evaluator assigns expansion mean scores of 4.94/5 for contextual coherence and 4.80/5 for dialogue naturalness. Expansion and reconstruction exhibit broadly similar interaction profiles; expansion's turn, overlap-event, backchannel, and interruption rates are 4.6%, 8.0%, 13.2%, and 16.0% lower, respectively. We evaluate data properties only; downstream gains in full-duplex model training remain for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。