arXiv:2501.06276cs.SDcs.CL2025-01被引 5

通过提示词控制情感与强度,让语音合成更自然有表现力。

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

  • 用提示词引导情感和强度,实现多说话人控制。
  • 结合大语言模型调节语调,保持内容不变。
  • 适合需要个性化语音表达的场景,如虚拟助手。

语音合成已从统计方法发展到深度神经网络架构,涌现出多种接近人类发音模式的文本转语音(TTS)模型。然而,捕捉语音中的情感与风格等细微差别仍具挑战。为此,我们提出一种基于提示的情感控制方法。该架构在多说话人背景下实现情感与强度的联合调控,并利用大语言模型(LLMs)操纵语音韵律,同时保留语言内容。通过嵌入情感线索、调节强度层级,并以提示词引导韵律变化,所提方法使合成语音具备类人表达力与多样性。最后,我们通过系统性探索上述控制机制的有效性验证了该方法的可行性。

原文摘要 · Abstract (English)

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion and style in speech synthesis is challenging. To address this challenge, we introduce an approach centered on prompt-based emotion control. The proposed architecture incorporates emotion and intensity control across multi-speakers. Furthermore, we leverage large language models (LLMs) to manipulate speech prosody while preserving linguistic content. Using embedding emotional cues, regulating intensity levels, and guiding prosodic variations with prompts, our approach infuses synthesized speech with human-like expressiveness and variability. Lastly, we demonstrate the effectiveness of our approach through a systematic exploration of the control mechanisms mentioned above.

语音合成情感控制提示词多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。