构建情感丰富且上下文详细的语音标注语料库,提升语音合成的情绪控制能力。
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
- 用生成模型自动提取并标注情感丰富的语音片段
- 通过自然语言描述增强情感粒度,减少人工标注成本
- 为情绪可控语音合成提供可扩展的高质量数据基础
文本到语音(TTS)技术虽已显著提升语音质量,接近目标说话人的音色与语调,但人类情感表达的复杂性仍使控制细微情绪差异成为巨大挑战。现有情感语音数据库多采用过于简化的标签体系,难以涵盖广泛情绪状态,限制了情绪合成效果。近期研究尝试使用自然语言描述情感,但成本高昂且情感深度不足。本文提出一种新方法:通过生成模型系统化提取情感丰富的语音段,并以详细自然语言描述进行标注。该方法提升了数据库的情感粒度,大幅降低对昂贵人工标注的依赖,实现高阶语言模型驱动的数据增强。最终构建的丰富语料库为开发更细腻、动态的情绪可控TTS系统提供了可扩展、经济可行的基础。
原文摘要 · Abstract (English)
Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional expression, the development of TTS systems capable of controlling subtle emotional differences remains a formidable challenge. Existing emotional speech databases often suffer from overly simplistic labelling schemes that fail to capture a wide range of emotional states, thus limiting the effectiveness of emotion synthesis in TTS applications. To this end, recent efforts have focussed on building databases that use natural language annotations to describe speech emotions. However, these approaches are costly and require more emotional depth to train robust systems. In this paper, we propose a novel process aimed at building databases by systematically extracting emotion-rich speech segments and annotating them with detailed natural language descriptions through a generative model. This approach enhances the emotional granularity of the database and significantly reduces the reliance on costly manual annotations by automatically augmenting the data with high-level language models. The resulting rich database provides a scalable and economically viable solution for developing a more nuanced and dynamic basis for developing emotionally controlled TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。