0.4B模型实现30语言零样本语音克隆,无需文本提示
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

- 用IPA统一表征+两阶段训练,无需对齐文本即可克隆语音
- 在420小时多语言数据上训练,支持30种语言零样本克隆
- 模型仅0.4B参数却媲美百亿级模型,适合快速部署
本文提出X-Voice,一个0.4B参数的多语言零样本语音克隆模型,可克隆任意语音并使用户能说30种语言。该模型基于420K小时的多语言语料库训练,采用国际音标(IPA)作为统一表示。为避免依赖带复杂预处理的提示文本,设计了两阶段训练:第一阶段通过标准条件流匹配训练得到X-Voice$_{\text{s1}}$,并用其合成10,000小时语音提示;第二阶段在掩码提示文本的音频对上微调,获得可实现零样本语音克隆的X-Voice$_{\text{s2}}$。架构上扩展F5-TTS,引入双层语言标识注入,并解耦调度无分类器引导以支持多语言语音合成。主观与客观评估显示,X-Voice优于现有基于流匹配的多语言系统如LEMAS-TTS,且零样本跨语言克隆能力媲美千亿级模型如Qwen3-TTS。所有资源已开源,促进研究透明与社区发展。
原文摘要 · Abstract (English)
In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voice$_{\text{s1}}$ through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voice$_{\text{s2}}$, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。