零样本生成未见情绪语音,用提示词+大模型实现情感混合控制
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
- 通过提示词学习融合大模型上下文知识,实现情绪可控的语音合成
- 在零样本条件下成功生成未训练过的情感组合语音,支持多情绪权重调节
- 适合需要灵活情感表达的对话系统、虚拟角色等场景
现有表情语音合成系统主要针对有限类别情绪建模,而人类对话远超这些预定义情绪,因此探索更丰富的情感语音生成对实现自然交互至关重要。本文提出一种新型零样本未见情绪语音合成方法(PUE),通过情感引导的提示学习实现未见情绪语音生成。PUE采用LLM-TTS架构训练,确保类别情绪相关提示与语音情感的一致性,使模型可量化每个语句中的不同情绪权重。推理时,通过灵活调整情绪比例并利用大模型上下文知识,实现混合情感语音生成,支持不同情感风格的定量表达。所提PUE在零样本设置下成功实现未见情绪的表达语音合成。
原文摘要 · Abstract (English)
Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech generation for more natural interactions. To bridge this gap, this paper proposes a novel prompt-unseen-emotion (PUE) approach to generate unseen emotional speech via emotion-guided prompt learning. PUE is trained utilizing an LLM-TTS architecture to ensure emotional consistency between categorical emotion-relevant prompts and emotional speech, allowing the model to quantitatively capture different emotion weightings per utterance. During inference, mixed emotional speech can be generated by flexibly adjusting emotion proportions and leveraging LLM contextual knowledge, enabling the model to quantify different emotional styles. Our proposed PUE successfully facilitates expressive speech synthesis of unseen emotions in a zero-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。