arXiv:2604.00688cs.CLeess.AS2026-04被引 22

OmniVoice实现600多种语言零样本语音合成,性能领先。

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models

  • 用扩散语言模型结构直接将文本转为多码本声学标记
  • 在58.1万小时开源数据上训练,覆盖超600种语言
  • 适合需要多语言语音生成的开发者和研究者

我们提出OmniVoice,一个可扩展至超过600种语言的大规模多语言零样本语音合成(TTS)模型。其核心是一种新型的扩散语言模型风格离散非自回归(NAR)架构。与传统两阶段(文本→语义→声学)复杂流程中性能受限的离散NAR模型不同,OmniVoice直接将文本映射到多码本声学标记。这一简化方法得益于两项关键技术:(1) 全码本随机掩码策略,实现高效训练;(2) 从预训练大语言模型初始化,保障高可懂性。通过利用总计58.1万小时、完全来自开源数据的多语言数据集,OmniVoice实现了迄今为止最广的语言覆盖,并在中文、英文及多样多语言基准上达到领先性能。代码与预训练模型已公开于https://github.com/k2-fsa/OmniVoice。

原文摘要 · Abstract (English)

We present OmniVoice, a massively multilingual zero-shot text-to-speech (TTS) model that scales to over 600 languages. At its core is a novel diffusion language model-style discrete non-autoregressive (NAR) architecture. Unlike conventional discrete NAR models that suffer from performance bottlenecks in complex two-stage (text-to-semantic-to-acoustic) pipelines, OmniVoice directly maps text to multi-codebook acoustic tokens. This simplified approach is facilitated by two key technical innovations: (1) a full-codebook random masking strategy for efficient training, and (2) initialization from a pre-trained LLM to ensure superior intelligibility. By leveraging a 581k-hour multilingual dataset curated entirely from open-source data, OmniVoice achieves the broadest language coverage to date and delivers state-of-the-art performance across Chinese, English, and diverse multilingual benchmarks. Our code and pre-trained models are publicly available at https://github.com/k2-fsa/OmniVoice.

语音合成多语言扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。