arXiv:2602.14664cs.SD2026-02

反向语音文本训练让端到端语音合成更自然

Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text Directions

  • 用反向文本-反向语音训练模型,探索人类发音约束的影响
  • 反向训练的模型在音质、可懂度和自然度上均优于传统方法
  • 适合研究语音生成机制或提升语音合成质量的研究者

端到端文本到语音(e2e-TTS)系统通过数据学习将文本与语音特征关联。人类语音涉及发音器官平滑过渡,但受解剖结构限制,某些发音状态难以实现或转换。本文实验研究这些生理约束对e2e-TTS训练的影响。采用两种架构:自回归的Tacotron-2和非自回归的VITS-TTS。构建三类系统:(a) 正向文本-正向语音(常规e2e-TTS)、(b) 反向文本-反向语音(r-e2e-TTS)、(c) 反向文本-正向语音(rtfs-e2e-TTS)。结果表明,e2e-TTS系统为纯数据驱动。有趣的是,r-e2e-TTS生成的语音在音质、感知可懂度和自然度方面表现更优。

原文摘要 · Abstract (English)

An end-to-end (e2e) text-to-speech (TTS) system is a deep architecture that learns to associate a text string with acoustic speech patterns from a curated dataset. It is expected that all aspects associated with speech production, such as phone duration, speaker characteristics, and intonation among other things are captured in the trained TTS model to enable the synthesized speech to be natural and intelligible. Human speech is complex, involving smooth transitions between articulatory configurations (ACs). Due to anatomical constraints, some ACs are challenging to mimic or transition between. In this paper, we experimentally study if the constraints imposed by human anatomy have an implication on training an e2e-TTS systems. We experiment with two e2e-TTS architectures, namely, Tacotron-2 an autoregressive model and VITS-TTS a non-autoregressive model. In this study, we build TTS systems using (a) forward text, forward speech (conventional, e2e-TTS), (b) reverse text, reverse speech (r-e2e-TTS), and (c) reverse text, forward speech (rtfs-e2e-TTS). Experiments demonstrate that e2e-TTS systems are purely data-driven. Interestingly, the generated speech by r-e2e-TTS systems exhibits better fidelity, better perceptual intelligibility, and better naturalness

语音合成端到端发音约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。