arXiv:2508.17623cs.CLeess.AS2025-08中稿 · at被引 4

构建对话系统情感推理评估基准,提升人机交互自然度。

EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems

  • 用文本转语音生成多样情绪数据,解决情绪语音样本稀缺问题。
  • 提出跨轮次情感推理得分,量化多轮对话中情绪变化一致性。
  • 评测7个系统发现情感不一致问题,助力构建更自适应对话模型。

语音情感在人机交互中至关重要,影响参与度与情境感知沟通。尽管对话系统近年取得进展,但全面评估情感推理能力的系统仍缺失。为此,我们提出EMO-Reasoning,一个用于评估对话系统情感连贯性的基准。该基准利用文本转语音生成的精选数据集,模拟多种情绪状态,克服了真实情绪语音数据匮乏的问题。我们进一步提出跨轮次情感推理得分,以评估多轮对话中的情绪演变。通过连续、分类和感知指标对七个对话系统进行评估,结果表明该框架能有效检测情感不一致,为改进现有系统提供洞见。通过发布系统性评估基准,我们旨在推动更具自然性与适应性的情感感知语音对话建模发展。

原文摘要 · Abstract (English)

Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reasoning is still lacking. To address this, we introduce EMO-Reasoning, a benchmark for assessing emotional coherence in dialogue systems. It leverages a curated dataset generated via text-to-speech to simulate diverse emotional states, overcoming the scarcity of emotional speech data. We further propose the Cross-turn Emotion Reasoning Score to assess the emotion transitions in multi-turn dialogues. Evaluating seven dialogue systems through continuous, categorical, and perceptual metrics, we show that our framework effectively detects emotional inconsistencies, providing insights for improving current dialogue systems. By releasing a systematic evaluation benchmark, we aim to advance emotion-aware spoken dialogue modeling toward more natural and adaptive interactions.

情感推理对话系统语音评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。