用自研对齐器提升语音合成时长精度,让发音更自然
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
- 先训练对齐器生成精准时长标签,替代外部工具
- 时长标注准确率提升,词错误率降低16%
- 适合追求高自然度语音合成的研究者与开发者
近期文本到语音(TTS)系统如FastSpeech和StyleSpeech显著提升了语音生成质量。然而,这些模型通常依赖外部工具(如Montreal Forced Aligner)生成时长,耗时且灵活性差。尽管时长对语调自然性和可懂性至关重要,其重要性常被低估。为此,我们提出一种新的对齐器引导训练范式,在训练TTS模型前先训练对齐器,以获得更精确的时长标注,减少对外部工具的依赖并提升对齐准确性。我们进一步研究了梅尔谱图、MFCCs及潜在特征等不同声学特征对TTS性能的影响。实验结果表明,对齐器引导的时长标注可使词错误率最高降低16%,显著改善音素与声调对齐效果。该方法有效优化了TTS系统,使其生成更自然、可懂的语音。
原文摘要 · Abstract (English)
Recent advancements in text-to-speech (TTS) systems, such as FastSpeech and StyleSpeech, have significantly improved speech generation quality. However, these models often rely on duration generated by external tools like the Montreal Forced Aligner, which can be time-consuming and lack flexibility. The importance of accurate duration is often underestimated, despite their crucial role in achieving natural prosody and intelligibility. To address these limitations, we propose a novel Aligner-Guided Training Paradigm that prioritizes accurate duration labelling by training an aligner before the TTS model. This approach reduces dependence on external tools and enhances alignment accuracy. We further explore the impact of different acoustic features, including Mel-Spectrograms, MFCCs, and latent features, on TTS model performance. Our experimental results show that aligner-guided duration labelling can achieve up to a 16\% improvement in word error rate and significantly enhance phoneme and tone alignment. These findings highlight the effectiveness of our approach in optimizing TTS systems for more natural and intelligible speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。