将语音编码器改至音素级,实现可解释的韵律解耦表征。
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
- 在音素级离散潜空间中建模韵律,分离发音与说话人信息。
- 潜空间主成分对应音高和能量,具有可解释性。
- 适合需要精细控制语音韵律的研究者或语音合成应用。
现有语音韵律建模方法多依赖连续潜空间中的全局风格表示来编码和迁移参考语音属性。然而,基于残差向量量化(RVQ)的神经编码器已展现出巨大潜力。本文研究了此类RVQ-VAE模型在音素级上的韵律建模能力,通过将编码器和解码器条件化于语言表示,并引入全局说话人嵌入,以解耦音素与说话人信息。通过主观实验和客观指标的系统评估,结果表明该方法获得的音素级离散潜表示具有高度解耦性,能捕捉细粒度且稳健可迁移的韵律信息。潜空间呈现出可解释结构,其主成分分别对应音高和能量。
原文摘要 · Abstract (English)
Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which are based on Residual Vector Quantization (RVQ) already shows great potential offering distinct advantages. We investigate the prosody modeling capabilities of the discrete space of such an RVQ-VAE model, modifying it to operate on the phoneme-level. We condition both the encoder and decoder of the model on linguistic representations and apply a global speaker embedding in order to factor out both phonetic and speaker information. We conduct an extensive set of investigations based on subjective experiments and objective measures to show that the phoneme-level discrete latent representations obtained this way achieves a high degree of disentanglement, capturing fine-grained prosodic information that is robust and transferable. The latent space turns out to have interpretable structure with its principal components corresponding to pitch and energy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。