arXiv:2508.11535eess.AS2025-08被引 1

让语音情感转换更自然,通过重合成控制发音时长。

Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling

  • 用重合成的离散内容表示建模语音时长,可精准调控情绪相关语速。
  • 在MSP-Podcast数据集上显著提升情感表达力,低唤醒情绪更慢更长。
  • 无需平行数据即可实现可控语速,适合真实场景语音转换应用。

语音情感转换旨在改变输入语音所表达的情绪,同时保持语义内容和说话人身份。近期生成模型在调整基频、谱包络和能量等局部声学特征方面表现良好,但难以控制语音时长。为此,我们提出一种基于重合成的离散内容表示时长建模框架,可在不使用平行数据的情况下,使语音时长反映目标情绪,实现可控语速。实验结果表明,在真实场景数据集MSP-Podcast上,该框架显著提升了情感表现力;分析显示,低唤醒情绪对应较长时长与较慢语速,高唤醒情绪则产生较短、较快的语音。

原文摘要 · Abstract (English)

Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic properties such as fundamental frequency, spectral envelope and energy, but often lack the ability to control the duration of sounds. To address this, we propose a duration modeling framework using resynthesis-based discrete content representations, enabling modification of speech duration to reflect target emotions and achieve controllable speech rates without using parallel data. Experimental results reveal that the inclusion of the proposed duration modeling framework significantly enhances emotional expressiveness, in the in-the-wild MSP-Podcast dataset. Analyses show that low-arousal emotions correlate with longer durations and slower speech rates, while high-arousal emotions produce shorter, faster speech.

语音转换情感识别时长建模无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。