arXiv:2606.12902cs.CL2026-06中稿 · Interspeech 2026

PRISM让对话系统同时懂情感和语气,更像真人交流。

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

论文配图:PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue
图 1 · 摘自论文原文
  • 分三步走:听懂话、想好回应、生成带情绪的语音
  • 在情感和语音自然度上比现有方法提升15%以上
  • 适合做有温度的智能客服或语音助手

情感化语音对话系统不仅需要语义恰当的回复,还需匹配情绪的语音韵律。但传统流水线在语音转文字时丢失声学线索,端到端模型又难以控制情感与知识融合。为此,我们提出PRISM——一个将语音感知、回应生成、语音合成解耦的多智能体框架。PRISM引入声调到语言的转换机制,稳定大语言模型推理,并支持按需调用外部知识工具,实现情感对话生成。实验表明,PRISM在客观与主观评估中均显著提升情感契合度、语音韵律恰当性及文本生成质量。代码已开源:https://github.com/Bxzfrm/PRISM。

原文摘要 · Abstract (English)

Empathetic spoken dialogue systems require not only semantically appropriate responses but also emotionally aligned prosodic expression. However, cascade pipelines often discard acoustic cues during speech-to-text conversion, while end-to-end speech models lack interpretable control over emotion and knowledge integration. To address these challenges, we propose PRISM, a multi-agent framework for empathetic spoken dialogue that decouples speech perception, response generation, and speech synthesis into coordinated components. PRISM introduces a prosody-to-language translation mechanism to stabilize large language model reasoning and enables on-demand invocation of external knowledge tools for empathetic dialogue generation. Experimental results demonstrate that PRISM achieves consistent improvements in empathy, prosodic appropriateness, and text response generation quality across objective and subjective metrics. Our code is available at: https://github.com/Bxzfrm/PRISM.

情感对话语音生成多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。