arXiv:2604.11594eess.AScs.SD2026-04被引 5

基于真人对话的多轮情感智能评测基准,更真实地评估语音语言模型的情感理解能力。

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models

论文配图:HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
图 1 · 摘自论文原文
  • 用真实人类对话重构多轮情感追踪与因果推理为选择题,减少主观评分偏差。
  • 发现多数模型在多轮情感追踪和隐含因果推理上表现不佳。
  • 揭示模型在跨模态冲突中严重依赖文本、忽视语音信号的偏见问题。

评估语音语言模型(ALMs)的情感智能(EI)至关重要。然而,现有基准大多依赖合成语音,仅限单轮交互,且高度依赖开放式评分。本文提出HumDial-EIBench,一个全面的ALMs情感智能评测基准。该基准使用来自ICASSP 2026 HumDial Challenge的真实人类对话数据,将情感追踪与因果推理转化为带有对抗性干扰项的多项选择题,降低认知任务中的主观评分偏差;保留共情回复生成任务,并引入声学-语义冲突任务,评估模型在矛盾多模态信号下的鲁棒性。对八种ALMs的评估显示,多数模型在多轮情感追踪和隐含因果推理方面表现不佳;所有模型均表现出文本与声学共情解耦,且在跨模态冲突中存在严重的文本主导偏见。

原文摘要 · Abstract (English)

Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn interactions, and depend heavily on open-ended scoring. This paper proposes HumDial-EIBench, a comprehensive benchmark for evaluating ALMs' EI. Using real-recorded human dialogues from the ICASSP 2026 HumDial Challenge, it reformulates emotional tracking and causal reasoning into multiple-choice questions with adversarial distractors, mitigating subjective scoring bias for cognitive tasks. It retains the generation of empathetic responses and introduces an acoustic-semantic conflict task to assess robustness against contradictory multimodal signals. Evaluations of eight ALMs reveal that most models struggle with multi-turn emotional tracking and implicit causal reasoning. Furthermore, all models exhibit decoupled textual and acoustic empathy, alongside a severe text-dominance bias during cross-modal conflicts.

情感智能语音模型多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。