arXiv:2412.11795cs.CLcs.SD2024-12AAAI被引 5

无需标注数据,就能让语音合成更自然地处理停顿和语调。

ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis

  • 用流匹配框架建模语音韵律,自动学习停顿位置与语调变化。
  • 在复杂长句上生成更准确的停顿时长和符合人感知的语调模式。
  • 适合需要精细控制语音节奏与情感表达的研究与应用。

韵律包含文字之外的丰富信息,对语音可懂度至关重要。现有模型在分句与语调方面仍存在不足,合成长句时常遗漏或错置停顿,且语调不自然。本文提出ProsodyFM,一种基于流匹配(Flow-Matching)的韵律感知文语转换(TTS)模型,旨在提升分句与语调表现。ProsodyFM引入两个核心组件:句段断点编码器用于捕捉初始断点位置,再通过持续时间预测器灵活调整断点时长;终端语调编码器学习一组语调形状令牌,并结合新颖的音高处理器,实现对人类感知语调变化的稳健建模。该模型无需显式韵律标签即可学习广泛的断点时长与语调模式。实验表明,相较四种SOTA模型,ProsodyFM显著改善了分句与语调,从而提升整体可懂度。分布外测试显示,其在未见复杂句式与说话人上具有更强泛化能力。案例研究直观展示了其对分句与语调的精细可控性。

原文摘要 · Abstract (English)

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks when synthesizing long sentences with complex structures but also produce unnatural intonation. We propose ProsodyFM, a prosody-aware text-to-speech synthesis (TTS) model with a flow-matching (FM) backbone that aims to enhance the phrasing and intonation aspects of prosody. ProsodyFM introduces two key components: a Phrase Break Encoder to capture initial phrase break locations, followed by a Duration Predictor for the flexible adjustment of break durations; and a Terminal Intonation Encoder which learns a bank of intonation shape tokens combined with a novel Pitch Processor for more robust modeling of human-perceived intonation change. ProsodyFM is trained with no explicit prosodic labels and yet can uncover a broad spectrum of break durations and intonation patterns. Experimental results demonstrate that ProsodyFM can effectively improve the phrasing and intonation aspects of prosody, thereby enhancing the overall intelligibility compared to four state-of-the-art (SOTA) models. Out-of-distribution experiments show that this prosody improvement can further bring ProsodyFM superior generalizability for unseen complex sentences and speakers. Our case study intuitively illustrates the powerful and fine-grained controllability of ProsodyFM over phrasing and intonation.

语音合成韵律控制无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。