arXiv:2511.12347eess.AScs.CL2025-11EMNLP被引 11

一个模型搞定11种语言的语音合成与编辑,零样本切换语种。

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

  • 用大语言模型处理文本,不依赖音素,实现跨语言统一建模。
  • 在仅有限语种数据下仍保持高质量语音生成与编辑效果。
  • 适合需要多语言语音应用的开发者和研究者快速部署使用。

我们提出 VoiceCraft-X,一种自回归神经编解码语言模型,统一实现11种语言(英语、中文、韩语、日语、西班牙语、法语、德语、荷兰语、意大利语、葡萄牙语、波兰语)的多语言语音编辑与零样本文本转语音(TTS)合成。该模型采用 Qwen3 大语言模型进行无音素的跨语言文本处理,并引入新颖的分词重排序机制,将时间对齐的文本与语音标记整合为单一序列生成任务。模型能生成高质量、自然流畅的语音,可在同一框架内无缝创建新音频或编辑已有录音。在数据稀缺的语种上也表现出稳健性能,凸显了统一自回归方法在复杂真实多语言语音应用中的潜力。音频样例见 https://zhishengzheng.com/voicecraft-x/。

原文摘要 · Abstract (English)

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/.

语音合成多语言语音编辑自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。