arXiv:2508.09600cs.SD2025-08被引 28

让语音对话机器人更懂情绪,提升共情能力。

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue

  • 分三阶段训练,强化语音理解与共情生成的联动。
  • 在20万条语音对话数据上表现优于现有模型。
  • 适合资源有限场景下的情感化语音交互研究。

共情对语音对话系统实现自然交互至关重要,使机器能识别并响应年龄、性别、情绪等副语言线索。近年来,统一语音理解与生成的端到端语音语言模型提供了可行方案,但仍面临三大挑战:过度依赖大规模对话数据集、对共情关键的副语言线索提取不足,以及缺乏专门的共情数据集与评估框架。为此,我们提出OSUM-EChat,一个开源的端到端语音对话系统,旨在增强资源受限环境下的共情交互。该系统引入两项核心创新:(1)三阶段理解驱动的语音对话训练策略,将大语音理解模型能力拓展至对话任务;(2)语言-副语言双重思维机制,通过思维链融合副语言理解与对话生成,提升回应共情度。该方法减少对大规模数据的依赖,同时保持高质量共情互动。此外,我们构建了EChat-200K数据集(20万条语音到语音的共情对话)和EChat-eval评估基准,全面评测系统共情能力。实验表明,OSUM-EChat在共情响应性方面优于现有端到端语音对话模型。

原文摘要 · Abstract (English)

Empathy is crucial in enabling natural interactions within spoken dialogue systems, allowing machines to recognize and respond appropriately to paralinguistic cues such as age, gender, and emotion. Recent advancements in end-to-end speech language models, which unify speech understanding and generation, provide promising solutions. However, several challenges persist, including an over-reliance on large-scale dialogue datasets, insufficient extraction of paralinguistic cues vital for conveying empathy, and the lack of empathy-specific datasets and evaluation frameworks. To address these issues, we introduce OSUM-EChat, an open-source, end-to-end spoken dialogue system designed to enhance empathetic interactions, particularly in resource-limited settings. OSUM-EChat introduces two key innovations: (1) a three-stage understanding-driven spoken dialogue training strategy that extends the capabilities of a large speech understanding model to spoken dialogue tasks, and (2) a linguistic-paralinguistic dual thinking mechanism that integrates paralinguistic understanding through a chain of thought with dialogue generation, enabling the system to produce more empathetic responses. This approach reduces reliance on large-scale dialogue datasets while maintaining high-quality empathetic interactions. Additionally, we introduce the EChat-200K dataset, a rich corpus of empathetic speech-to-speech dialogues, and the EChat-eval benchmark, a comprehensive framework for evaluating the empathetic capabilities of dialogue systems. Experimental results demonstrate that OSUM-EChat outperforms end-to-end spoken dialogue models regarding empathetic responsiveness, validating its effectiveness.

语音对话共情生成副语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。