首个基于理论的情感智能评估框架,让语音模型真正懂情绪。
EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

- 构建四维情感认知理论框架,覆盖感知、理解、运用和管理情绪
- 领先模型仅达52.6%准确率,人类基准远超此水平
- 提出新评估模型EmoS,接近人类83.8%表现,适用于真实对话场景
尽管指令遵循和听觉理解取得进展,但语音语言模型(SLMs)的情感智能(EI)评估仍局限于基础的副语言感知,缺乏系统性、理论驱动的认知框架。我们提出EmoSBench,首个基于四分支理论模型的全面情感智能评估基准,涵盖感知、理解、运用和管理情绪的十项子任务。初步评估显示显著差距:即使领先专有模型GPT-4o-Audio也仅达52.6%,远低于人类基准。为弥合这一差距,我们开发了经监督微调(SFT)与组相对策略优化(GRPO)训练的EmoS专用评估模型。为此,我们构建了EmoDialogue双语数据集,通过精细标注的情绪梯度响应对提供监督信号。同时引入融合陡峭指数精度奖励(SEAR)与论据一致性奖励(RFR)的机制,确保精确序数评分与有效推理。实验表明,EmoS达到83.8%准确率,接近人类水平。在真实、无约束语音交互上的评估验证了其强泛化能力,为提升情感智能对话系统奠定了基础框架。
原文摘要 · Abstract (English)
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。