arXiv:2602.10716eess.AScs.CL2026-02中稿 · IEEE ASRU 2025被引 1

用语音情绪细节提升AI共情能力,效果显著优于纯文本模型。

RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance

  • 融合维度化情绪嵌入与辅助学习,增强语音情感理解
  • 在多个数据集上共情评分提升超14%,情绪探索能力翻倍以上
  • 适合研究情感交互、语音生成与人机共情的学者与开发者

随着生成式AI的发展,人机交互中的共情能力愈发重要。以往研究多关注情绪反射,而情绪探索——实现更深层次互动的关键——却常被忽视。现有大模型依赖文本,难以捕捉丰富的情绪细微差别。为此,我们提出RE-LLM,一种集成维度化情绪嵌入与辅助学习的语音-大模型。实验表明,在三个数据集上,其共情指标均取得统计显著提升:在ESD上,相对文本基线和语音基线,情绪反应得分分别提高14.79%和6.76%;在IEMOCAP上,探索得分提升35.42%和3.91%;在ESD上提升139.28%和9.83%;在MSP-PODCAST上提升60.95%和22.64%。此外,在语音情感识别中,未加权准确率分别提升5.4%(IEMOCAP)、2.3%(ESD)和6.9%(MSP-PODCAST)。结果凸显了RE-LLM在情绪理解深度与共情响应生成上的优势。

原文摘要 · Abstract (English)

With generative AI advancing, empathy in human-AI interaction is essential. While prior work focuses on emotional reflection, emotional exploration, key to deeper engagement, remains overlooked. Existing LLMs rely on text which captures limited emotion nuances. To address this, we propose RE-LLM, a speech-LLM integrating dimensional emotion embeddings and auxiliary learning. Experiments show statistically significant gains in empathy metrics across three datasets. RE-LLM relatively improves the Emotional Reaction score by 14.79% and 6.76% compared to text-only and speech-LLM baselines on ESD. Notably, it raises the Exploration score by 35.42% and 3.91% on IEMOCAP, 139.28% and 9.83% on ESD, and 60.95% and 22.64% on MSP-PODCAST. It also boosts unweighted accuracy by 5.4% on IEMOCAP, 2.3% on ESD, and 6.9% on MSP-PODCAST in speech emotion recognition. These results highlight the enriched emotional understanding and improved empathetic response generation of RE-LLM.

语音生成共情模型情绪识别大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。