用大模型实现自然语言自由控制语音情绪,让合成语音更传情达意。
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
- 利用大语言模型实现自然语言描述的情绪精准控制
- 在英语数据集上达到当前最佳效果,中文也表现优异
- 自建40小时高质量情绪语音数据集,支持真实场景评估
人类语音不仅是信息传递,更是情感交流与人际连接。尽管语音合成(TTS)模型取得显著进展,但在控制生成语音的情感表达方面仍面临挑战。本文提出EmoVoice,一种基于大语言模型(LLM)的可控情感语音合成模型,支持通过自由风格自然语言指令精细调控情感。采用受思维链(CoT)和跨模态链(CoM)启发的并行输出设计,同时生成音素与音频标记,提升内容一致性。我们还构建了高质量40小时英语情绪语音数据集EmoVoice-DB,包含丰富表达与自然语言描述的情感标签。EmoVoice仅使用合成数据,在英语EmoVoice-DB测试集上达到领先性能;在中文Secap测试集上也表现优秀。研究还评估了现有情感评价指标与人类感知偏好的一致性,并探索使用GPT-4o-audio与Gemini等前沿多模态大模型进行情感语音评测。代码、模型权重及演示样本已开源。
原文摘要 · Abstract (English)
Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and chain-of-modality (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints, and demo samples are available at https://github.com/yanghaha0908/EmoVoice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。