arXiv:2501.04904cs.CLcs.SD2025-01中稿 · ICASSP 2025被引 4

用大模型联合识别情绪与上下文,让语音合成更自然。

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis

  • 用情绪感知编码器让大模型理解语音情绪
  • 在有限数据下仍能精准生成情感匹配的对话语音
  • 适合需要情感化语音合成的智能客服、虚拟助手

近年来,对话式语音合成(CSS)对自然语音生成的需求日益增长,需考虑对话上下文。为此,我们提出JELLY,一种新型CSS框架,通过微调大语言模型(LLM)并引入多个部分LoRA模块,实现情绪识别与上下文推理的联合建模。我们设计了情绪感知的Q-former编码器,使LLM能够感知语音中的情绪,该编码器利用情感语音数据集进行训练,以对齐语音情绪与文本语义。随后,整个模型在对话语音数据上进一步微调,从而推断情感上下文,生成符合对话情境的情绪化语音。实验结果表明,JELLY在情感上下文建模方面表现优异,生成的语音与对话自然契合,同时有效缓解了情感对话语音数据稀缺的问题。

原文摘要 · Abstract (English)

Recently, there has been a growing demand for conversational speech synthesis (CSS) that generates more natural speech by considering the conversational context. To address this, we introduce JELLY, a novel CSS framework that integrates emotion recognition and context reasoning for generating appropriate speech in conversation by fine-tuning a large language model (LLM) with multiple partial LoRA modules. We propose an Emotion-aware Q-former encoder, which enables the LLM to perceive emotions in speech. The encoder is trained to align speech emotions with text, utilizing datasets of emotional speech. The entire model is then fine-tuned with conversational speech data to infer emotional context for generating emotionally appropriate speech in conversation. Our experimental results demonstrate that JELLY excels in emotional context modeling, synthesizing speech that naturally aligns with conversation, while mitigating the scarcity of emotional conversational speech datasets.

语音合成情绪识别大模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。