通过声学与语义联合建模,提升语音情感理解与合成效果。
Acoustic and Semantic Modeling of Emotion in Spoken Language
- 提出声学与语义协同预训练方法,学习更具情感表征力的语音表示。
- 在对话场景中实现更优的情感识别性能,且风格迁移语音可用于数据增强。
- 无需人工标注文本即可实现大规模情感感知文本建模,适合语音情感系统开发。
情感在人类交流中起核心作用,影响信任、参与度与社会互动。随着大语言模型融入日常生活,让人工智能可靠地理解与生成人类情感仍是一大挑战。尽管情感表达具有多模态特性,本论文聚焦于语音中的情感表达,研究如何联合建模声学与语义信息以推动情感理解与语音合成的发展。第一部分提出通过预训练进行情感感知表示学习,引入声学与语义监督策略,使模型更好捕捉语音中的情感线索;同时提出一种语音驱动的监督预训练框架,可在无手动标注文本语料的情况下实现大规模情感感知文本建模。第二部分针对对话场景中的情感识别,设计结合跨模态注意力与专家混合融合的分层架构,整合多轮对话中的声学与语义信息。最后,提出一种无需并行数据的文本无关语音到语音情感风格迁移框架,可实现可控情感转换,同时保留说话人身份与语言内容。实验表明,该方法显著提升情感迁移效果,且风格迁移语音可用于数据增强,进而提高情感识别性能。
原文摘要 · Abstract (English)
Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to reliably understand and generate human emotions remains an important challenge. While emotional expression is inherently multimodal, this thesis focuses on emotions conveyed through spoken language and investigates how acoustic and semantic information can be jointly modeled to advance both emotion understanding and emotion synthesis from speech. The first part of the thesis studies emotion-aware representation learning through pre-training. We propose strategies that incorporate acoustic and semantic supervision to learn representations that better capture affective cues in speech. A speech-driven supervised pre-training framework is also introduced to enable large-scale emotion-aware text modeling without requiring manually annotated text corpora. The second part addresses emotion recognition in conversational settings. Hierarchical architectures combining cross-modal attention and mixture-of-experts fusion are developed to integrate acoustic and semantic information across conversational turns. Finally, the thesis introduces a textless and non-parallel speech-to-speech framework for emotion style transfer that enables controllable emotional transformations while preserving speaker identity and linguistic content. The results demonstrate improved emotion transfer and show that style-transferred speech can be used for data augmentation to improve emotion recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。