发现对话情绪识别中标签模糊问题,提出按语境评估情绪可能性的新方法。
Exposing Weaknesses in Emotion Recognition in Conversations

- 用大模型零样本分析对话情绪,发现错误集中在否定、感叹等表达上。
- 人工重标注显示仅35%案例达成一致,情绪标签普遍存在歧义。
- 提出情绪独立评估框架,突破单标签局限,适合研究标注可靠性者。
对话中的情绪识别(ERC)旨在识别多轮对话中说话人的情绪。准确的情绪识别可支持共情对话代理、心理健康支持及教育技术等多种应用。尽管近期方法多依赖特定任务微调,此类模型可能利用数据集特有线索。当前ERC的核心假设是每句话可赋予单一明确情绪标签,但这一假设极少被质疑。我们采用大语言模型(LLMs)在零样本设置下,结合前序对话作为上下文研究ERC。结果显示,总体指标掩盖了系统性失败:错误集中于包含否定、感叹和感叹词的语句,且该现象在所有评估模型中均一致,暗示基准缺陷而非模型自身弱点。一项由四位人工标注员参与的受控重标注研究支持此发现:仅有35%案例达成强一致性,中性语句主导高一致性实例,而多数情感类别处于低一致性区间。这些结果表明,许多看似模型错误实为标注本身的模糊性所致。因此,标准单标签评估不足以反映真实情况。为解决此问题,我们提出一种‘大模型作裁判’框架,根据语境独立评估每个情绪的合理性,而非强制单标签决策。
原文摘要 · Abstract (English)
Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。