arXiv:2505.17589cs.SDcs.AI2025-05被引 212

CosyVoice 3实现零样本多语言语音生成,支持海量数据与更大模型规模。

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

  • 用多任务监督训练新语音分词器,提升语调自然度。
  • 扩展至100万小时数据,覆盖9种语言和18种方言。
  • 模型参数增至15亿,支持更复杂场景的语音合成。

在前期工作中,我们提出了可扩展的流式语音合成模型CosyVoice 2,融合大语言模型(LLM)与块感知流匹配(FM)模型,实现低延迟双流语音合成与人类水平音质。然而,CosyVoice 2在语言覆盖、领域多样性、数据量、文本格式及后训练技术方面仍存局限。本文提出CosyVoice 3,一种面向真实场景的零样本多语言语音生成模型,在内容一致性、说话人相似性和韵律自然度上超越前代。主要改进包括:1)通过自动语音识别、语音情感识别、语言识别、音频事件检测和说话人分析等多任务监督训练,构建新语音分词器,显著提升韵律自然度;2)提出可微分奖励模型用于后训练,适用于各类基于LLM的语音合成模型;3)数据量从十万小时扩展至一百万小时,涵盖9种语言和18种中文方言,覆盖多种领域与文本格式;4)模型参数从0.5亿增至15亿,凭借更强容量在多语言基准测试中表现更优。这些进展显著推动了真实场景语音合成的发展。演示可访问 https://funaudiollm.github.io/cosyvoice3。

原文摘要 · Abstract (English)

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.

语音合成多语言大模型后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。