arXiv:2606.11681cs.CLcs.SD2026-06中稿 · Interspeech 2026, …

用罗马化统一书写系统,让语音合成支持495种语言

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

论文配图:UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction
图 1 · 摘自论文原文
  • 将多种文字转为罗马音统一表示,突破百语言限制
  • 在495种语言上表现优于现有模型,低资源语言也有效
  • 适合需要多语言语音合成的开发者和研究者

我们提出UR-BERT,一种基于罗马化转写的文字到语音(TTS)编码器,用于大规模多语言TTS系统。传统基于字符到音素(G2P)的方法受限于可靠G2P资源,仅能支持约100种语言。相比之下,UR-BERT通过将多样书写系统统一为共享的罗马化表示,扩展至495种语言。为进一步提升语音保真度与文本-语音对齐,我们在训练中引入语音标记预测目标,使编码器以数据高效方式学习语音感知的音素表示。实验表明,基于UR-BERT构建的TTS系统在多种语言和资源条件下均持续优于近期文本编码器基线,并展现出对未见语言的强大泛化能力。

原文摘要 · Abstract (English)

We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable G2P resources. In contrast, UR-BERT scales to 495 languages by unifying diverse writing systems into a shared Romanization representation. To further enhance phonetic fidelity and text-speech alignment, we introduce a speech token prediction objective during training, which encourages the encoder to learn speech-aware phonetic representations in a data-efficient manner. Experiments show that TTS systems built on UR-BERT consistently outperform recent text encoder baselines across a wide range of languages and resource conditions, and demonstrate strong generalization to unseen languages.

多语言合成语音生成文本编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。