用ASR引导的在线优化,让少语料语言也能生成自然语音。
Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
- 基于GRPO框架,用无配对数据优化语音合成模型。
- 在少语料语言上生成可懂且说话人一致的语音,效果优于单纯微调。
- 适合资源匮乏语言的语音合成,也提升高资源语言表现。
为低资源语言开发高质量文本到语音(TTS)系统面临配对文本-语音数据稀缺的挑战。相比之下,由于大规模多语言预训练,这些语言的自动语音识别(ASR)模型往往更易获取。本文提出一种基于组相对策略优化(GRPO)的框架,用于适配自回归多语言TTS模型至新语言。首先,使用国际音标(IPA)token训练一个语言无关的TTS基线模型;其次,在有限的配对数据上微调该模型以捕捉目标语言的韵律特征;最后,利用仅含无配对文本与说话人提示的GRPO方法,通过预训练的ASR、说话人验证和音频质量评估模型构建多目标奖励进行优化。实验表明,该流程在低资源语言中生成了可懂且说话人一致的语音,显著优于单纯微调。此外,该GRPO框架在高资源语言上也表现更优,超越离线对齐方法如直接偏好优化(DPO),在可懂性、说话人相似性和音频质量方面均有提升。
原文摘要 · Abstract (English)
Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech recognition (ASR) models for such languages are often more accessible, owing to large-scale multilingual pre-training efforts. We propose a framework based on Group Relative Policy Optimization (GRPO) to adapt an autoregressive, multilingual TTS model to new languages. Our method first establishes a language-agnostic foundation for TTS synthesis by training a multilingual baseline with International Phonetic Alphabet (IPA) tokens. Next, we fine-tune this model on limited paired data of the new languages to capture the target language's prosodic features. Finally, we apply GRPO to optimize the model using only unpaired text and speaker prompts, guided by a multi-objective reward from pretrained ASR, speaker verification, and audio quality estimation models. Experiments demonstrate that this pipeline produces intelligible and speaker-consistent speech in low-resource languages, substantially outperforming fine-tuning alone. Furthermore, our GRPO-based framework also improves TTS performance in high-resource languages, surpassing offline alignment methods such as Direct Preference Optimization (DPO) yielding superior intelligibility, speaker similarity, and audio quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。