实现高质量零样本多人对话语音合成,保持情感连贯性与音色一致性。
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

- 采用分层训练策略,从单人朗读逐步过渡到真实对话数据。
- 在双语评测集上超越所有开源基线的丰富度与层次感评分。
- 支持1-4人对话,适用于需要自然互动语音的应用场景。
零样本文本转语音(TTS)在单说话人合成方面已取得显著进展,但富有表现力的长篇多人对话合成仍具挑战。常见做法是用单人朗读TTS模型逐句合成后拼接,导致推理开销大、声学不一致、对话连贯性差及情感断续。近期对话TTS系统虽有所改进,但仍难以同时兼顾表现力连贯性、可控说话人切换和单人朗读质量。本文提出SwanData-Speech与SwanVoice:SwanData-Speech基于真实音频构建朗读与对话语料库,使用Swan强制对齐器进行停顿感知的词级对齐,并以RobustMegaTTS3处理发音困难案例;在此数据基础上,SwanVoice是面向1–4名说话人的零样本TTS模型,结合25 Hz VAE、带停顿感知符号与拼音替换的原始文本条件,以及含说话人回合条件的流匹配DiT。训练从单人语音开始,经混合数据、真实对话数据,再通过DiffusionNFT后训练,引入音素级与说话人相似性奖励。在SwanBench-Speech评测中,SwanVoice在朗读与对话设置下均优于所有评估的开源基线,在丰富度与层次感得分上领先,内容准确性仍是主要短板。音频演示可访问 https://swanaigc.github.io//#swanvoice。
原文摘要 · Abstract (English)
Zero-shot text-to-speech (TTS) has improved substantially for single-speaker synthesis, yet expressive long-form multi-speaker dialogue remains difficult. A common workaround is to synthesize each turn with a monologue TTS model and stitch the outputs together. This adds inference cost and often breaks acoustic consistency, conversational coherence, and affective continuity across turns. Recent dialogue TTS systems have begun to address this setting, but they still struggle to keep expressive coherence, controllable speaker switching, and monologue quality at the same time. We present SwanData-Speech and SwanVoice. SwanData-Speech builds monologue and dialogue corpora from in-the-wild audio, using Swan Forced Aligner for pause-aware word-level alignment and RobustMegaTTS3 for pronunciation-hard cases. Built on these data, SwanVoice is a zero-shot TTS model for 1--4 speakers, combining a 25 Hz VAE, raw-text conditioning with pause-aware symbols and pinyin substitution, and a flow-matching DiT with speaker-turn conditioning. Training starts from monologue speech, moves through mixed and real dialogue data, and then uses DiffusionNFT post-training with phone-level and speaker-similarity rewards. On SwanBench-Speech, SwanVoice obtains higher richness and hierarchy scores than all evaluated open-source baselines in both monologue and dialogue settings, while content accuracy remains the main limitation. Audio demos are available at https://swanaigc.github.io//#swanvoice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。