MOSS-TTSD可生成多角色、长时长的自然对话语音,支持零样本音色克隆。
MOSS-TTSD: Text to Spoken Dialogue Generation
- 基于增强的长上下文建模,直接从带发言者标签的剧本生成对话语音
- 支持60分钟单次合成、5方对话及零样本音色克隆,覆盖中英文等主流语言
- 提出新评估框架TTSD-eval,无需说话人分离即可客观衡量对话质量
口语对话生成对播客、动态解说和娱乐内容至关重要,但相比单轮文本转语音(TTS),面临准确换言、跨轮次声学一致性及长篇稳定性等挑战。现有模型因缺乏对话上下文建模能力而表现不佳。为此,我们提出MOSS-TTSD,一个面向多语言、多角色表达性对话语音的合成模型。该模型具备增强的长上下文建模能力,可从带发言者标签的对话脚本生成长达60分钟的单次合成语音,支持最多5名说话人参与,且能通过短参考音频实现零样本音色克隆。模型支持英语、中文等主流语言,适用于多种长篇场景。为克服现有评估方法局限,我们提出基于强制对齐的客观评估框架TTSD-eval,无需依赖说话人分离工具即可衡量说话人归属准确率与相似度。客观与主观评测均表明,MOSS-TTSD在对话合成上超越强开源及专有基线。
原文摘要 · Abstract (English)
Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate turn-taking, cross-turn acoustic consistency, and long-form stability, which current models often fail to address due to a lack of dialogue context modeling. To bridge this gap, we present MOSS-TTSD, a spoken dialogue synthesis model designed for expressive, multi-party conversational speech across multiple languages. With enhanced long-context modeling, MOSS-TTSD generates long-form spoken conversations from dialogue scripts with explicit speaker tags, supporting up to 60 minutes of single-pass synthesis, multi-party dialogue with up to 5 speakers, and zero-shot voice cloning from a short reference audio clip. The model supports various mainstream languages, including English and Chinese, and is adapted to several long-form scenarios. Additionally, to address limitations of existing evaluation methods, we propose TTSD-eval, an objective evaluation framework based on forced alignment that measures speaker attribution accuracy and speaker similarity without relying on speaker diarization tools. Both objective and subjective evaluation results show that MOSS-TTSD surpasses strong open-source and proprietary baselines in dialogue synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。