arXiv:2409.18512cs.SDcs.AI2024-09被引 1

通过两阶段提示选择,提升零样本语音合成的情感强度与说话人一致性。

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS

  • 先用韵律特征和LLM评估候选提示,再在具体模型中测试发音错误率、说话人与情感相似度。
  • 动态阶段根据输入文本匹配最契合的提示,提升生成适配性。
  • 适合追求自然情感表达与稳定说话人特征的语音合成研究者。

近期语音合成进展使得基于大语言模型(LLM)的系统能够在无需训练的情况下,通过输入提示实现内容、语调、说话人身份和情感的可控生成。然而,现有提示选择方法往往无法确保提示包含足够稳定的说话人身份线索和恰当的情感强度指示,这对富有表现力的语音合成至关重要。为此,我们提出一种专为富有表现力语音合成设计的两阶段提示选择策略。静态阶段(合成前)利用基于音高韵律特征、感知音频质量及由LLM评估的文本-情感一致性得分,对候选提示进行初步筛选;进一步通过特定TTS模型评估字符错误率、说话人相似度和合成语音与提示语音之间的情感相似度。动态阶段(合成过程中)使用文本相似度模型,选择与当前输入文本最匹配的提示。实验结果表明,该策略能有效选出兼具高情感强度与强说话人一致性的提示,显著提升零样本语音合成的表现力与稳定性。音频样例与代码将公开于 https://whyrrrrun.github.io/ExpPro.github.io/。

原文摘要 · Abstract (English)

Recent advancements in speech synthesis have enabled large language model (LLM)-based systems to perform zero-shot generation with controllable content, timbre, speaker identity, and emotion through input prompts. As a result, these models heavily rely on prompt design to guide the generation process. However, existing prompt selection methods often fail to ensure that prompts contain sufficiently stable speaker identity cues and appropriate emotional intensity indicators, which are crucial for expressive speech synthesis. To address this challenge, we propose a two-stage prompt selection strategy specifically designed for expressive speech synthesis. In the static stage (before synthesis), we first evaluate prompt candidates using pitch-based prosodic features, perceptual audio quality, and text-emotion coherence scores evaluated by an LLM. We further assess the candidates under a specific TTS model by measuring character error rate, speaker similarity, and emotional similarity between the synthesized and prompt speech. In the dynamic stage (during synthesis), we use a textual similarity model to select the prompt that is most aligned with the current input text. Experimental results demonstrate that our strategy effectively selects prompt to synthesize speech with both high-intensity emotional expression and robust speaker identity, leading to more expressive and stable zero-shot TTS performance. Audio samples and codes will be available at https://whyrrrrun.github.io/ExpPro.github.io/.

语音合成零样本情感表达提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。