arXiv:2508.18655cs.CLcs.SD2025-08被引 5

用小数据训练出能共情的语音对话模型,说话更懂人心。

Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models

  • 基于小规模数据构建情感理解与回应生成框架
  • 在20万条情感对话数据上实现4.41分语音质量与3.97分共情得分
  • 适合追求高情感表达但算力有限的语音助手研发

随着语音大语言模型的发展,用户现在可以通过语音直接与助手交互。然而,现有模型大多只将回复内容转为语音,未能充分捕捉用户语句中丰富的表情情绪,相同句子因语气不同含义各异。因此,情感理解对提升人机交互至关重要。目前大多数共情语音大模型依赖海量数据,计算成本高。关键挑战在于如何在数据有限、无需大规模预训练的情况下,构建能生成共情回应的模型。为此,我们提出 Emotion Omni,该模型可理解用户语音中的情感内容并生成共情回复。我们还构建了一个支持共情语音助手的20万条情感对话数据集。实验表明,Emotion Omni 在无需大规模预训练的前提下,仍具备可比的指令遵循能力,同时在语音质量(UTMOS: 4.41)和共情表现(Emotion GPT Score: 3.97)上优于现有模型。结果证实其在语音保真度与情感表达性上的双重提升。演示地址:https://w311411.github.io/omni_demo/

原文摘要 · Abstract (English)

With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich emotional cues in user queries, where the same sentence may convey different meanings depending on the expression. Emotional understanding is thus essential for improving human-machine interaction. Most empathetic speech LLMs rely on massive datasets, demanding high computational cost. A key challenge is to build models that generate empathetic responses with limited data and without large-scale training. To this end, we propose Emotion Omni, a model that understands emotional content in user speech and generates empathetic responses. We further developed a data pipeline to construct a 200k emotional dialogue dataset supporting empathetic speech assistants. Experiments show that Emotion Omni achieves comparable instruction-following ability without large-scale pretraining, while surpassing existing models in speech quality (UTMOS:4.41) and empathy (Emotion GPT Score: 3.97). These results confirm its improvements in both speech fidelity and emotional expressiveness. Demos are available at https://w311411.github.io/omni_demo/.

共情对话语音生成小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。