首个让人类与大模型同台比拼情感支持对话的评估框架
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
- 设计统一评测体系,让真人和大模型在相同对话中互相对比
- 顶尖模型在共情感知上接近甚至超过人类平均水平
- 适合研究人机情感交互、评估对话模型真实社交能力的学者
情感支持对话依赖语言流畅性之外的能力,如情绪识别、语气调整及应对抵触、沮丧等情境。尽管大语言模型发展迅速,我们仍缺乏清晰方法来比较其人际互动能力与人类的差异。本文提出HEART,首个直接在多轮情感支持对话中对比人类与大模型表现的框架。对每段对话历史,配对真人与模型回复,并通过盲评人类评分员与多模型评判器组成的集成系统进行评估。所有评估基于人际沟通科学的五维量表:人类契合度、共情响应、调谐性、共鸣度与任务遵循度。结果显示,若干前沿模型在感知共情与一致性上达到或超越人类平均表现;而人类在适应性重构、情绪命名与微妙语气变化方面仍具优势,尤其在对抗性对话中。人类与模型评判者在约80%的成对比较中达成一致,接近人与人之间的一致性水平,且评语聚焦相似维度。这表明支持质量的评价标准正趋于趋同。通过将人类与模型置于同一基准,HEART将情感支持对话重构为独立于通用推理与语言流利性的能力维度,为理解模型生成支持与人类社会判断的契合点、分歧点及其随模型规模增长的演化提供了统一实证基础。
原文摘要 · Abstract (English)
Supportive conversation depends on skills that go beyond language fluency, including reading emotions, adjusting tone, and navigating moments of resistance, frustration, or distress. Despite rapid progress in language models, we still lack a clear way to understand how their abilities in these interpersonal domains compare to those of humans. We introduce HEART, the first-ever framework that directly compares humans and LLMs on the same multi-turn emotional-support conversations. For each dialogue history, we pair human and model responses and evaluate them through blinded human raters and an ensemble of LLM-as-judge evaluators. All assessments follow a rubric grounded in interpersonal communication science across five dimensions: Human Alignment, Empathic Responsiveness, Attunement, Resonance, and Task-Following. HEART uncovers striking behavioral patterns. Several frontier models approach or surpass the average human responses in perceived empathy and consistency. At the same time, humans maintain advantages in adaptive reframing, tension-naming, and nuanced tone shifts, particularly in adversarial turns. Human and LLM-as-judge preferences align on about 80 percent of pairwise comparisons, matching inter-human agreement, and their written rationales emphasize similar HEART dimensions. This pattern suggests an emerging convergence in the criteria used to assess supportive quality. By placing humans and models on equal footing, HEART reframes supportive dialogue as a distinct capability axis, separable from general reasoning or linguistic fluency. It provides a unified empirical foundation for understanding where model-generated support aligns with human social judgment, where it diverges, and how affective conversational competence scales with model size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。