arXiv:2506.10150cs.CLcs.HC2025-06被引 27

LLMs在共情沟通判断上表现可靠,优于普通人,接近专家水平。

When Large Language Models are Reliable for Judging Empathic Communication

  • 用心理学、NLP等4个框架评估200段真实对话中的共情表达
  • LLMs在4个框架中均接近专家一致性,高于普通人群标注者
  • 适合用于情感敏感场景的共情能力验证与监督

大型语言模型(LLMs)在生成共情回应方面表现优异,但其对共情沟通细微差别的判断可靠性如何?我们通过对比专家、众包工作者和LLMs在四个来自心理学、自然语言处理与传播学的评估框架下对200段真实对话的标注结果,展开研究。基于3,150份专家标注、2,844份众包标注和3,150份LLM标注,评估三类标注者间的评分一致性。结果显示,专家一致性强,但不同框架子维度因清晰度、复杂性和主观性差异而异。我们证明,专家一致性比标准分类指标更能有效衡量LLM性能。在所有四个框架中,LLMs始终接近专家基准,并显著超过众包标注者的可靠性。这些发现表明,当在特定任务上经过适当基准验证时,LLMs可支持情感敏感应用(如对话伴侣)中的透明度与监督。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at generating empathic responses in text-based conversations. But, how reliably do they judge the nuances of empathic communication? We investigate this question by comparing how experts, crowdworkers, and LLMs annotate empathic communication across four evaluative frameworks drawn from psychology, natural language processing, and communications applied to 200 real-world conversations where one speaker shares a personal problem and the other offers support. Drawing on 3,150 expert annotations, 2,844 crowd annotations, and 3,150 LLM annotations, we assess inter-rater reliability between these three annotator groups. We find that expert agreement is high but varies across the frameworks' sub-components depending on their clarity, complexity, and subjectivity. We show that expert agreement offers a more informative benchmark for contextualizing LLM performance than standard classification metrics. Across all four frameworks, LLMs consistently approach this expert level benchmark and exceed the reliability of crowdworkers. These results demonstrate how LLMs, when validated on specific tasks with appropriate benchmarks, can support transparency and oversight in emotionally sensitive applications including their use as conversational companions.

共情识别大模型评估人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。