arXiv:2506.00885cs.SDcs.AI2025-06NeurIPS被引 15

无需自回归生成,实现零样本多说话人对话合成

CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching

  • 采用流匹配模型直接从多路转录文本生成语音频谱
  • 在语音质量与说话人一致性上超越现有方法
  • 支持重叠语音和精确时间控制,适合真实场景

生成自然流畅的多说话人对话对播客创作、虚拟助手和多媒体内容生成至关重要。现有系统难以保持说话人一致性、建模重叠语音,且合成效率低。本文提出 CoVoMix2,一种完全非自回归的零样本多说话人对话生成框架。该模型基于流匹配机制,直接从多路转录文本生成梅尔频谱,避免依赖中间标记表示。为更好捕捉真实对话动态,我们引入转录级说话人解耦、句级对齐与提示级随机掩码策略。实验表明,CoVoMix2 在语音质量、说话人一致性和推理速度上均优于 MoonCast、Sesame 等强基线。特别地,该方法无需提示转录即可运行,并支持重叠语音与精确时间控制,展现出对真实语音生成场景的强大泛化能力。

原文摘要 · Abstract (English)

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently. In this paper, we introduce CoVoMix2, a fully non-autoregressive framework for zero-shot multi-talker dialogue generation. CoVoMix2 directly predicts mel-spectrograms from multi-stream transcriptions using a flow-matching-based generative model, eliminating the reliance on intermediate token representations. To better capture realistic conversational dynamics, we propose transcription-level speaker disentanglement, sentence-level alignment, and prompt-level random masking strategies. Our approach achieves state-of-the-art performance, outperforming strong baselines like MoonCast and Sesame in speech quality, speaker consistency, and inference speed. Notably, CoVoMix2 operates without requiring transcriptions for the prompt and supports controllable dialogue generation, including overlapping speech and precise timing control, demonstrating strong generalizability to real-world speech generation scenarios.

对话生成非自回归流匹配多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。