arXiv:2505.12597cs.SDeess.AS2025-05ACL被引 6

让语音合成更懂情绪,模仿人类对话理解流程。

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

  • 分三阶段模拟人类对话理解:情绪识别、语义压缩、共情合成
  • 在三个数据集上表现优于现有方法,情感表达更自然
  • 适合研究情感语音合成与人机共情交互的开发者

对话式语音合成(CSS)旨在使合成语音与用户-代理交互的情感和风格语境一致,实现共情。当前生成式CSS模型因情感感知不足和冗余离散语音编码而存在可解释性缺陷。为此,我们提出Chain-Talker,一个三阶段框架,模拟人类认知过程:情绪理解从对话历史中提取上下文相关的表情描述;语义理解通过序列化预测生成紧凑的语义码;共情渲染则融合两者生成富有表现力的语音。为支持情感建模,我们构建了CSS-EmCap,一个基于大语言模型的自动化管道,用于生成精确的对话语音情感描述。在三个基准数据集上的实验表明,Chain-Talker生成的语音更具表现力和共情性,且CSS-EmCap有效提升了情感建模可靠性。代码与演示见:https://github.com/AI-S2-Lab/Chain-Talker。

原文摘要 · Abstract (English)

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. To address the above issues, we present Chain-Talker, a three-stage framework mimicking human cognition: Emotion Understanding derives context-aware emotion descriptors from dialogue history; Semantic Understanding generates compact semantic codes via serialized prediction; and Empathetic Rendering synthesizes expressive speech by integrating both components. To support emotion modeling, we develop CSS-EmCap, an LLM-driven automated pipeline for generating precise conversational speech emotion captions. Experiments on three benchmark datasets demonstrate that Chain-Talker produces more expressive and empathetic speech than existing methods, with CSS-EmCap contributing to reliable emotion modeling. The code and demos are available at: https://github.com/AI-S2-Lab/Chain-Talker.

语音合成情感建模共情交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。