用自然语言精准控制语音情感,生成更自然的带情绪语音。
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
- 通过对比学习与扩散模型,根据文本描述生成情感声学特征。
- 仅用文本条件即可调节音高、抖动和音量等语音属性。
- 在保持语音质量的同时实现细粒度情感控制,适合交互式语音系统。
当前情感语音合成(TTS)系统虽能生成清晰的情感语音,但对情感表现的精细控制仍具挑战。本文提出ParaEVITS框架,利用自然语言的组合性增强情感渲染的可控性。该框架基于受ParaCLAP启发的文本-音频编码器,通过对比语言-音频预训练(CLAP)模型,使扩散模型能够根据文本情感描述生成情感嵌入。训练阶段先使用参考音频进行音频编码器预训练,再微调扩散模型以处理ParaCLAP的文本编码器输入。推理时,仅通过文本条件即可调控音高、抖动和响度等语音属性。实验表明,ParaEVITS在不牺牲语音质量的前提下有效控制情感渲染,语音示例已公开。
原文摘要 · Abstract (English)
While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce ParaEVITS, a novel emotional TTS framework that leverages the compositionality of natural language to enhance control over emotional rendering. By incorporating a text-audio encoder inspired by ParaCLAP, a contrastive language-audio pretraining (CLAP) model for computational paralinguistics, the diffusion model is trained to generate emotional embeddings based on textual emotional style descriptions. Our framework first trains on reference audio using the audio encoder, then fine-tunes a diffusion model to process textual inputs from ParaCLAP's text encoder. During inference, speech attributes such as pitch, jitter, and loudness are manipulated using only textual conditioning. Our experiments demonstrate that ParaEVITS effectively control emotion rendering without compromising speech quality. Speech demos are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。