arXiv:2510.22758cs.CL2025-10中稿 · ICLR被引 8

首个评估语音模型共情能力的多层级综合基准,测试其理解语义与声调情绪的结合能力。

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

  • 设计跨任务连贯的四阶段评测流程,模拟人类共情对话的认知过程。
  • 12个先进语音模型在高表达力声调识别上表现不佳,共情响应质量受限。
  • 适用于评估语音助手、心理陪伴机器人等需要情感智能的场景。

语音语言模型(SLMs)在口语理解方面取得显著进展,但尚不清楚它们能否充分感知非词汇性声音线索并基于情感与上下文因素做出共情回应。现有基准多孤立评估语言、声学、推理或对话能力,忽视了这些技能整合对类人情感对话的关键作用。我们提出EchoMind,首个关联式、多层级评测基准,通过顺序、上下文连贯的任务模拟共情对话认知过程:语音内容理解、声学线索感知、综合推理与响应生成。所有任务使用相同且语义中性的脚本,避免显性情感或上下文提示,通过控制语音风格变化来测试表达方式对模型的影响。EchoMind基于涵盖3个粗粒度与12个细粒度维度的共情框架,包含39个声学属性,采用客观与主观双重指标评估。测试12个先进SLMs发现,即使顶尖模型在高表达力声调识别上仍表现不足,限制了共情响应质量。分析显示,模型在指令遵循、对自然语音变异的鲁棒性以及有效利用声学线索进行共情方面存在持续弱点。结果表明,需构建能融合语言内容与多样化声学线索的语音模型,才能实现真正共情的对话能力。

原文摘要 · Abstract (English)

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks typically evaluate linguistic, acoustic, reasoning, or dialogue abilities in isolation, overlooking the integration of these skills that is crucial for human-like, emotionally intelligent conversation. We present EchoMind, the first interrelated, multi-level benchmark that simulates the cognitive process of empathetic dialogue through sequential, context-linked tasks: spoken-content understanding, vocal-cue perception, integrated reasoning, and response generation. All tasks share identical and semantically neutral scripts that are free of explicit emotional or contextual cues, and controlled variations in vocal style are used to test the effect of delivery independent of the transcript. EchoMind is grounded in an empathy-oriented framework spanning 3 coarse and 12 fine-grained dimensions, encompassing 39 vocal attributes, and evaluated using both objective and subjective metrics. Testing 12 advanced SLMs reveals that even state-of-the-art models struggle with high-expressive vocal cues, limiting empathetic response quality. Analyses of prompt strength, speech source, and ideal vocal cue recognition reveal persistent weaknesses in instruction-following, resilience to natural speech variability, and effective use of vocal cues for empathy. These results underscore the need for SLMs that integrate linguistic content with diverse vocal cues to achieve truly empathetic conversational ability.

语音模型共情对话多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。