评测语音对话模型的社交情商,发现现有模型仍依赖文本、易陷入安全陷阱。
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

- 构建包含2265段对话的多轮评估框架,基于情感商数理论设计评分体系。
- 实测显示当前模型在语音情感理解上存在认知偏差,尤其在长对话中遗忘上下文。
- 适合研究对话系统情商、人机交互与多模态情感计算的研究者使用。
随着多模态对话系统越来越多地参与口语交互,其处理副语言社交线索的能力已成为自然人机沟通的关键瓶颈。现有机器情感智能评估仅通过孤立文本或被动声学感知进行,忽略了主动多轮对话所需的跨模态推理。我们提出 extsc{SpeechEQ},一个用于评估语音-语言模型(SLMs)社会语言推理能力的综合性框架。该框架包含基于 EQ-i 2.0 理论构建的 2,265 段对话数据集,覆盖 15 个情商子维度,并引入受人类情商测评启发的口语情商(SEQ)评分机制。实验表明,现有语音情感识别与端到端语音-语言模型在理解与运用语音副语言线索方面存在明显局限。尽管端到端架构优于级联系统, extsc{SpeechEQ} 揭示当前多模态模型仍受限于文本依赖的“模态捷径”、对齐引发的“安全陷阱”以及“情境失忆”,凸显实现真正情感智能的障碍。该基准可访问 https://huggingface.co/datasets/SpeechEQ/SpeechEQ,演示页为 https://binomial14.github.io/speecheq-demo/
原文摘要 · Abstract (English)
As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critical bottleneck for natural human-AI communication. However, existing evaluations of machine emotional intelligence assess reasoning exclusively through isolated text or passive acoustic perception, overlooking the complex cross-modal reasoning required for active, multi-turn dialogue. We introduce \textsc{SpeechEQ}, a comprehensive framework designed to evaluate the sociolinguistic reasoning of Speech-Language Models (SLMs). The framework includes a validated dataset of 2,265 dialogues across 15 Emotional Quotient (EQ) subscales grounded in EQ-i 2.0 theory, along with a multi-turn evaluation protocol measured by our proposed Spoken EQ (SEQ) score inspired by human EQ assessments. Experiments show limitations in how both existing Speech Emotion Recognition and end-to-end Speech-Language Models understand and apply paralinguistic cues through speech. While end-to-end architectures outperform cascaded systems, \textsc{SpeechEQ} reveals that current multimodal models remain bottlenecked by a text-reliant ``modality shortcut,'' an alignment-induced ``safety trap,'' and ``contextual amnesia,'' highlighting the barriers to truly emotionally aware AI. Our benchmark can be accessed at https://huggingface.co/datasets/SpeechEQ/SpeechEQ and demo page at https://binomial14.github.io/speecheq-demo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。