arXiv:2605.21739cs.AI2026-05

构建真实对话场景下的情感智能评估基准,揭示大模型情感响应的多维度能力。

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence

论文配图:AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
图 1 · 摘自论文原文
  • 基于200段真实人机多轮对话与逐轮情绪标注,构建情感智能评估数据集。
  • 11个模型在情绪识别、行为分类、偏好预测等任务上排名独立,说明情感智能可分解。
  • 偏好对齐和回应质量判断比情绪标签准确率更能区分模型表现,适合真实场景应用。

情感智能(EI)是感知、理解并恰当回应他人情绪状态的能力,在人类交流中至关重要,随着大语言模型(LLMs)在日常对话中扮演角色,其评估愈发重要。现有情感智能评测依赖合成提示、单轮对话或第三方标注,无法直接衡量模型在真实对话中推断并回应参与者情绪状态的能力。我们提出AttuneBench,基于200段真实的多轮人机对话,参与者与匿名化LLMs对话,并对每一轮提供情绪状态、模型行为及期望回应的标注。在11个被测模型中,情绪识别、行为分类、偏好预测与回应质量评价的排名高度独立,表明情感智能可分解为可分离的能力模块。偏好对齐和回应质量判断显著优于情绪标签准确率,具有更强的模型区分能力。结果表明,真正的情感智能需预测特定用户在上下文中的期望回应,这一差异在聚合评分中可能被掩盖,且单轮或合成格式无法捕捉跨轮次动态。AttuneBench为评估各项能力提供了框架,可用于诊断模型在情感敏感对话中的优势与失效模式。

原文摘要 · Abstract (English)

Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly important to assess as LLMs assume conversational roles in everyday life. Existing EI benchmarks rely on synthetic prompts, single-turn cases, or third-party annotation. These approaches do not directly measure how models infer and respond to a participant's emotional state over the course of a real conversation. We introduce AttuneBench, a benchmark grounded in 200 genuine multi-turn human-model conversations in which participants conversed with anonymized LLMs and provided turn-by-turn annotations of their emotional state, the model's behavior, and their preferred responses. Across 11 evaluated models, we find that model rankings on emotion recognition, behavioral classification, preference prediction, and judged response quality are largely independent, indicating that emotionally intelligent behavior decomposes into separable capabilities. Preference alignment and response-quality judgments are substantially more model-discriminating than emotion-label accuracy. These results indicate that emotionally intelligent behavior requires predicting what kind of response a specific user wants in context, a distinction that aggregate scoring can obscure and that single-turn or synthetic formats cannot directly capture across turns. AttuneBench provides a framework for assessing each of these capabilities and for diagnosing model-specific strengths and failure modes in emotionally salient conversation.

情感智能对话评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。