arXiv:2511.00850eess.AScs.AI2025-11被引 15

首个评估语音对话模型情感智能的多轮互动基准

MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models

  • 设计分层任务结构,涵盖情绪理解与支持应用
  • 3200+样本验证,发现模型在复杂互动中表现不足
  • 适合研究对话系统情感智能的学者和开发者

语音对话模型发展迅速,但其在真实多轮交互对话中的能力仍缺乏深入评估,现有基准多集中于单轮对话。我们提出Multi-Bench,首个专为评估语音对话模型多轮交互能力及情感智能而设计的基准。该基准采用分层结构,包含基础赛道(情绪理解与推理)与进阶赛道(情绪支持与应用),涵盖五个精心设计的任务,共约3200个样本,覆盖从情绪识别到复杂推理与互动对话的全链条场景,并提供可复现的评估框架。我们在八大子集上对六种代表性语音对话模型进行了评测。结果显示,当前模型在基础理解任务上表现良好,但在高级多轮交互与推理相关任务中仍有提升空间,尤其在情绪觉察与实际应用方面表现不足。

原文摘要 · Abstract (English)

Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on single-turn exchanges. We introduce Multi-Bench, the first benchmark explicitly designed to evaluate SDMs in multi-turn interactive dialogue with an emphasis on emotional intelligence. Multi-Bench employs a hierarchical structure with a basic track for emotion understanding and reasoning and an advanced track for emotion support and application. It comprises five carefully designed tasks and about 3.2K samples, ranging from emotion recognition to complex reasoning and interactive dialogue, supported by a reproducible evaluation framework. We evaluate six representative SDMs on eight subsets of Multi-Bench. Results show that while current SDMs achieve good performance on basic understanding tasks, they still have room for improvement in advanced multi-turn interactive dialogue and reasoning-related tasks, particularly in emotion awareness and application.

对话系统情感智能多轮对话评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。