arXiv:2512.19090cs.SDeess.AS2025-12被引 8

JoyVoice实现八人长期对话的自然语音合成,支持多语言零样本克隆。

JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis

  • 统一端到端架构直接用自回归隐表示驱动扩散模型。
  • 12.5Hz低比特率多模态分词器提升语义与声学建模能力。
  • 支持多语言长对话生成,语音自然度与韵律连续性领先。

大型语音生成模型正从单说话人短句合成转向多说话人长对话生成。现有模型多局限于二人轮流交互。为此,我们提出JoyVoice,一种面向灵活、无边界合成的拟人化基础模型,可支持最多八名说话人。不同于传统级联系统,JoyVoice采用统一的端到端Transformer-DiT架构,直接利用自回归隐表示作为扩散模型输入,实现整体端到端优化。我们进一步提出低比特率(12.5 Hz)的多模态分词器(MM-Tokenizer),结合多任务语义与最小均方误差损失,有效建模语义与声学信息。模型还通过大规模数据扰动实现鲁棒文本前端处理。实验表明,JoyVoice在中、英、日、韩多语言生成及零样本语音克隆任务上达到顶尖水平,在Seed-TTS-Eval基准和多说话人长对话语音克隆任务中表现优异,显著提升长篇语音的韵律连续性、多说话人对话的节奏丰富性、副语言自然度与可懂度。

原文摘要 · Abstract (English)

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice

语音合成多说话人长对话零样本克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。