arXiv:2606.26534cs.SDcs.AI2026-06中稿 · Interspeech 2026

用强化学习在推理时快速适配,让语音合成更像陌生口音和特殊场景的声音。

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

论文配图:VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation
图 1 · 摘自论文原文
  • 通过强化学习在推理阶段优化语音风格,无需重新训练。
  • 在罕见口音和混杂语音上表现显著优于现有方法。
  • 适合需要快速个性化语音合成的场景,如虚拟助手、有声书生成。

近期零样本语音合成(TTS)实现了高保真与富有表现力的语音生成,但在模仿不常见场景中的说话风格(如对话干扰、方言)时仍存在困难。此外,微调预训练模型需要大量高质量数据,限制了快速个性化。本文提出 VoiceTTA,一种基于强化学习的推理时适应(TTA)方法,用于提升预训练零样本 TTS 模型的语音模仿能力。VoiceTTA 引入两种基于基频(F0)和能量系数变异率差异的风格奖励,结合说话人相似性与可懂度(使用预训练 Whisper 模型计算的词错误率,WER),并通过分组相对偏好优化(GRPO)在流匹配模型中优化可学习前缀,在推理时实现自适应调整。大量实验表明,该方法在罕见语音提示下性能显著提升,超越当前最优基线。音频样例见 https://voicetta.pages.dev/

原文摘要 · Abstract (English)

Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at https://voicetta.pages.dev/

语音合成强化学习零样本推理时适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。