arXiv:2507.00227eess.AScs.AI2025-07中稿 · Interspeech 2025被引 1

用随机方法提升语音合成中韵律的自然度与可控性

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis

  • 采用归一化流等随机建模技术生成韵律参数
  • 生成韵律与真人发音在主观评测中表现相当
  • 可通过调节采样温度实现更精细的控制

尽管生成模型近年发展迅速,文本到语音合成中生成富有表现力的韵律仍具挑战,尤其在显式建模音高、能量和时长等参数以增强可解释性与可控性的系统中。本文研究了归一化流、条件流匹配和修正流等随机方法在该任务中的有效性,并与传统确定性基线及真实人类发音进行对比。通过广泛的主观与客观评估,结果表明随机方法能有效捕捉人类语音固有的变异性,生成的韵律自然度达到与真人相当水平。此外,该方法支持通过调节采样温度实现额外的可控性选项。

原文摘要 · Abstract (English)

While generative methods have progressed rapidly in recent years, generating expressive prosody for an utterance remains a challenging task in text-to-speech synthesis. This is particularly true for systems that model prosody explicitly through parameters such as pitch, energy, and duration, which is commonly done for the sake of interpretability and controllability. In this work, we investigate the effectiveness of stochastic methods for this task, including Normalizing Flows, Conditional Flow Matching, and Rectified Flows. We compare these methods to a traditional deterministic baseline, as well as to real human realizations. Our extensive subjective and objective evaluations demonstrate that stochastic methods produce natural prosody on par with human speakers by capturing the variability inherent in human speech. Further, they open up additional controllability options by allowing the sampling temperature to be tuned.

语音合成韵律建模随机方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。