arXiv:2507.12015cs.SDeess.AS2025-07中稿 · INTERSPEECH 2025被引 5

让语音合成同时自然表达情绪和强调重点

EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis

  • 用伪标签和方差特征实现弱监督强调控制
  • 情绪变化时仍能保持强调位置清晰稳定
  • 适合需要精准情感与强调的语音应用

近年来,情感文本转语音(TTS)和强调可控语音合成取得了显著进展,但二者之间的交互仍待深入。本文提出强调与情绪融合的TTS框架(EME-TTS),旨在解决两个关键问题:(1)如何有效利用强调提升情感语音的表现力;(2)如何在不同情绪下维持目标强调的感知清晰度与稳定性。EME-TTS采用弱监督学习方法,结合强调伪标签与基于方差的强调特征,并引入强调感知增强(EPE)模块,强化情绪信号与强调位置间的交互。实验表明,结合大语言模型预测强调位置时,EME-TTS可生成更自然的情感语音,且在不同情绪间保持稳定、可区分的目标强调。合成样例已在线公开。

原文摘要 · Abstract (English)

In recent years, emotional Text-to-Speech (TTS) synthesis and emphasis-controllable speech synthesis have advanced significantly. However, their interaction remains underexplored. We propose Emphasis Meets Emotion TTS (EME-TTS), a novel framework designed to address two key research questions: (1) how to effectively utilize emphasis to enhance the expressiveness of emotional speech, and (2) how to maintain the perceptual clarity and stability of target emphasis across different emotions. EME-TTS employs weakly supervised learning with emphasis pseudo-labels and variance-based emphasis features. Additionally, the proposed Emphasis Perception Enhancement (EPE) block enhances the interaction between emotional signals and emphasis positions. Experimental results show that EME-TTS, when combined with large language models for emphasis position prediction, enables more natural emotional speech synthesis while preserving stable and distinguishable target emphasis across emotions. Synthesized samples are available on-line.

语音合成情感语音强调控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。