通过分阶段预训练提升语音合成的韵律感知能力
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS

- 先用掩码语言建模,再用混合音素对比学习训练双流编码器
- 混合音素对比阶段显著提升语音合成质量,尤其在语义清晰度和说话人相似度上
- 单独优化音素一致性会损害发音区分度,影响生成效果
我们研究了基于扩散模型的语音合成中韵律建模的多阶段预训练方法。采用说话人条件下的双流编码器,先通过掩码语言建模训练,再使用混合音素批次进行类似SigLIP的跨模态对比学习,并额外研究了同音素精修阶段。在Grad-TTS和潜在扩散语音合成系统上评估了文本-音频检索性能及下游合成效果。两阶段课程(掩码语言建模 + 混合音素对比学习)在可懂度、说话人相似性和感知评价指标上均取得最佳合成质量。尽管同音素精修提升了韵律检索性能,但降低了音素区分度并导致合成质量下降。结果表明,嵌入空间指标的提升未必带来生成性能改善,强调在语音合成预训练中需平衡音素区分度与韵律敏感性。
原文摘要 · Abstract (English)
We investigate multi-stage pretraining for prosody modeling in diffusion-based TTS. A speaker-conditioned dual-stream encoder is trained with masked language modeling followed by SigLIP-style cross-modal contrastive learning using mixed-phoneme batches, with an additional same-phoneme refinement stage studied separately. We evaluate intrinsic text-audio retrieval and downstream synthesis in Grad-TTS and a latent diffusion TTS system. The two-stage curriculum (MLM + mixed-phoneme contrastive learning) achieves the best overall synthesis quality in terms of intelligibility, speaker similarity, and perceptual measures. Although same-phoneme refinement improves prosodic retrieval, it reduces phoneme discrimination and degrades synthesis. These findings indicate that improvements in embedding-space metrics do not necessarily translate to better generative performance and highlight the need to balance phoneme discrimination and prosodic sensitivity in TTS pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。