arXiv:2606.25424eess.AScs.AI2026-06中稿 · INTERSPEECH 2026

提出可自适应调制周期性的新非线性,提升语音合成中情感语调的细腻表现力。

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS

论文配图:Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS
图 1 · 摘自论文原文
  • 设计自适应振荡非线性,通过线性旁路保持信号稳定
  • 在LJSpeech和情感语音数据集上均实现客观与主观评分提升
  • 适合需要精细控制语调变化的语音合成研究者

基于扩散模型的文语转换(TTS)在语音质量上已取得显著进步,但对富有表现力语音中的尖锐语调过渡和快速音高变化建模仍具挑战。现有扩散模型解码器常用周期性非线性如Snake激活函数捕捉谐波结构,但在处理突发的幅度与频率变化时适应性不足。本文研究了振荡归纳偏置在扩散模型解码器中的作用,提出一种自适应振荡非线性,可在保持信号稳定性的同时实现可控的周期调制。该方法构建的TTS系统称为OscillaTTS。在LJSpeech和情感语音数据集上的实验显示,该方法在客观与主观评价上均有持续提升,表明其更优地建模了富有表现力的语调动态。

原文摘要 · Abstract (English)

Diffusion-based text-to-speech (TTS) models have achieved significant improvements in speech quality. However, modeling sharp prosodic transitions and rapid pitch variations in expressive speech remains challenging. Existing diffusion-based TTS decoders commonly utilize periodic nonlinearities such as Snake activation function to capture harmonic structures, but this activation funcation provides limited adaptability when modeling abrupt amplitude and frequency variations. In this paper, we investigate the role of oscillatory inductive bias in diffusion-based TTS decoders and introduce an adaptive oscillatory nonlinearity that enables controllable periodic modulation while maintaining signal stability through a linear bypass component. We refer the resulting TTS system as OscillaTTS. Experiments on the LJSpeech and Emotional Speech Dataset show consistent improvements across objective and subjective evaluations, indicating improved modeling of expressive prosodic dynamics.

语音合成扩散模型语调建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。