用轻量方法精准控制多语种语音合成中日语的发音和声调。
UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
- 基于低秩适配,实现日语发音与声调的音素级调控。
- 零样本设置下保持语音自然度与说话人相似性不变。
- 适合需要精准语言发音控制的语音合成应用。
我们提出 UtterTune,一种基于大语言模型(LLM)构建的多语种文本到语音(TTS)系统的轻量级适配方法。该方法在不损害其他语言性能的前提下,提升了目标语言(本文为日语)的发音控制能力。尽管大语言模型架构已使 TTS 模型达到极高的自然度,但缺乏显式音素-音位(G2P)模块且直接处理最小编码文本(如字节对编码)时,准确建模音素到音位映射与韵律仍具挑战。UtterTune 利用低秩适配,在零样本设定下实现了对日语语音的音段发音与声调的音素级精细控制,同时维持了语音自然度与说话人相似性。客观与主观评估验证了其有效性。
原文摘要 · Abstract (English)
We propose UtterTune, a lightweight method for adapting a multilingual text-to-speech (TTS) system built on a large language model (LLM). It improves control of pronunciation in the target language while preserving performance in the others. Although LLM architectures have enabled TTS models to achieve remarkable naturalness, accurately modeling grapheme-to-phoneme (G2P) mapping and prosody remains challenging, especially when the model omits an explicit G2P module and directly processes minimally encoded text (e.g., byte-pair encoding). UtterTune leverages low-rank adaptation to enable the control of segmental pronunciation and pitch accent at the phoneme level for Japanese speech, the target language in this paper, while maintaining naturalness and speaker similarity in a zero-shot setting. Objective and subjective evaluations confirm its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。