通过显式建模情感上下文提升语音对话共情生成效果
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
- 先建模情感状态再编码语音,避免信息丢失
- 双层注意力捕捉单轮与全程情感动态,提升连贯性
- 适合需要高情感智能的语音客服、心理陪伴场景
共情语音对话不仅需理解语言内容,还需感知韵律、语调和情感强度等非语言信息。现有语音到语音的大模型多依赖语音识别转写或编码器提取隐表示,常导致情感信息弱化与多轮对话连贯性下降。为此,我们提出ES4R框架,核心在于在语音编码前显式建模结构化情感上下文,而非依赖编码器隐式学习或显式情绪标注。具体采用双层注意力机制,分别捕捉单轮情感状态与全程情感动态。所得情感表征通过语音引导的跨模态注意力与文本语义融合,生成共情回应。语音输出阶段,采用基于能量的策略选择与风格融合实现共情语音合成。在自动与人工评估中,ES4R持续优于强基线,并在不同大模型底座下保持鲁棒性。
原文摘要 · Abstract (English)
Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing speech-to-speech large language models either rely on ASR transcription or use encoders to extract latent representations, often weakening affective information and contextual coherence in multi-turn dialogues. To address this, we propose \textbf{ES4R}, a framework for speech-based empathetic response generation. Our core innovation lies in explicitly modeling structured affective context before speech encoding, rather than relying on implicit learning by the encoder or explicit emotion supervision. Specifically, we introduce a dual-level attention mechanism to capture turn-level affective states and dialogue-level affective dynamics. The resulting affective representations are then integrated with textual semantics through speech-guided cross-modal attention to generate empathetic responses. For speech output, we employ energy-based strategy selection and style fusion to achieve empathetic speech synthesis. ES4R consistently outperforms strong baselines in both automatic and human evaluations and remains robust across different LLM backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。