arXiv:2606.06928cs.SDeess.AS2026-06被引 17

2B参数多语言语音生成模型,支持30种语言和方言的可控合成。

VoxCPM2 Technical Report

论文配图:VoxCPM2 Technical Report
图 1 · 摘自论文原文
  • 统一框架整合30语言、9种方言及风格化克隆,全靠连续潜变量建模。
  • 音频重构速率48kHz,编码16kHz,实现高效隐式超分辨率。
  • 开源模型权重与工具,适合语音合成研究与工业落地使用。

我们提出VoxCPM2,一个完全开源的多语言可控语音生成基础模型,扩展了VoxCPM的分层扩散-自回归建模范式。VoxCPM2在三个关键维度取得进展:(i) 能力,统一支持30种语言、9种中文方言、自然语言语音设计、风格可控语音克隆以及高保真续写克隆;(ii) 质量,采用非对称AudioVAE,编码速率16 kHz,重构速率48 kHz,实现高效隐式超分辨率;(iii) 规模,模型参数达20亿,训练数据超过200万小时多语言语音。为在单一模型中支持多样能力,引入统一序列组织机制,通过同一输入模块的不同排列表达所有生成模式,实现单套参数与目标下的联合训练。VoxCPM2在公开的零样本和指令跟随语音合成基准上达到业界领先或竞争力表现。在内部30语言评估集上,平均字错误率(WER)为1.68%。结果表明,无需外部离散语音分词器的分层连续潜变量建模,可成为大规模多语言可控语音生成的有效且强大的基础。模型权重、微调代码与推理工具已按Apache 2.0许可开源,以推动社区研究与发展。

原文摘要 · Abstract (English)

We present VoxCPM2, a https://info.arxiv.org/help/prep#abstractsfully open-source multilingual and controllable speech generation foundation model that extends the hierarchical diffusion-autoregressive modeling paradigm of VoxCPM. VoxCPM2 advances the framework in three key dimensions: (i) capability, by unifying 30 languages, 9 Chinese dialects, natural-language voice design, style-controllable voice cloning, and high-fidelity continuation cloning within a single backbone; (ii) quality, through an asymmetric AudioVAE that encodes at 16 kHz and reconstructs at 48 kHz, enabling implicit super-resolution with high encoding efficiency; and (iii) scale, by jointly scaling the model to 2B parameters and the training data to over 2 million hours of multilingual speech. To support these diverse capabilities within one model, we introduce a unified sequence organization that expresses all generation modes through different arrangements of the same input building blocks, allowing joint training under a single set of parameters and objective. VoxCPM2 achieves state-of-the-art or competitive performance on public zero-shot and instruction-following TTS benchmarks. On our internal 30-language evaluation set, it attains an average WER of 1.68%. These results demonstrate that hierarchical continuous-latent modeling, without relying on any external discrete speech tokenizer, offers a viable and powerful foundation for large-scale multilingual and controllable speech generation. The model weights, fine-tuning code, and inference tools are publicly released under the Apache 2.0 license to foster community research and development.

语音生成多语言连续潜变量开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。