arXiv:2507.14988eess.AS2025-07AAAI被引 8

用强化学习优化语音合成中的时长预测,提升音质与多样性。

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

  • 通过强化学习优化时长预测模块,使用语者相似度和词错误率作奖励信号。
  • 在多个指标上超越先前系统,采样步数减半且音质不下降。
  • 适合关注语音合成质量与效率的研究者及开发者。

基于扩散模型的文本到语音(TTS)系统在零样本语音合成方面取得显著进展,但对感知指标的整体优化仍具挑战。以往的DMOSpeech工作直接优化生成组件的感知指标,但时长预测部分未被优化。本文提出DMOSpeech 2,通过强化学习将指标优化扩展至时长预测器。系统采用新颖的时长策略框架,结合组相对偏好优化(GRPO),以语者相似度和词错误率作为奖励信号。通过优化这一此前未优化的模块,DMOSpeech 2构建了更完整的指标优化合成流程。此外,本文引入教师引导采样,利用教师模型完成初始去噪步骤后切换至学生模型,显著提升输出多样性并保持效率。全面评估表明,该系统在所有指标上均优于先前方法,采样步数减少一半且无质量损失。这些进展标志着多组件指标优化语音合成的重要突破。音频样本、代码及预训练模型可于https://dmospeech2.github.io/获取。

原文摘要 · Abstract (English)

Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric optimization for speech generation components, but duration prediction remained unoptimized. This paper presents DMOSpeech 2, which extends metric optimization to the duration predictor through a reinforcement learning approach. The proposed system implements a novel duration policy framework using group relative preference optimization (GRPO) with speaker similarity and word error rate as reward signals. By optimizing this previously unoptimized component, DMOSpeech 2 creates a more complete metric-optimized synthesis pipeline. Additionally, this paper introduces teacher-guided sampling, a hybrid approach leveraging a teacher model for initial denoising steps before transitioning to the student model, significantly improving output diversity while maintaining efficiency. Comprehensive evaluations demonstrate superior performance across all metrics compared to previous systems, while reducing sampling steps by half without quality degradation. These advances represent a significant step toward speech synthesis systems with metric optimization across multiple components. The audio samples, code and pre-trained models are available at https://dmospeech2.github.io/.

语音合成扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。