评测大模型精准定位情绪表达片段的能力,推动情感理解更精细。
SEER: The Span-based Emotion Evidence Retrieval Benchmark
- 以句子片段为单位,识别具体表达情绪的文本段落。
- 在1200条真实语句上构建新标注数据集,覆盖单句与五句短文。
- 发现模型在长文本中表现下降,易误判中性内容为情绪表达。
我们提出SEER(基于片段的情绪证据检索基准),用于评估大语言模型识别具体情绪表达文本片段的能力。与传统情感分类任务不同,SEER聚焦于尚未充分研究的情绪证据检测——精确找出哪些短语传达了情绪。这种细粒度方法对共情对话和临床支持等应用至关重要,因需理解情绪如何被表达,而不仅是判断情绪类型。SEER包含两项任务:单句内情绪证据识别,以及连续五句话中的跨句证据识别。数据集新增了1200条真实语句的情感及情绪证据标注。我们评估了14个开源LLM,发现部分模型在单句输入下接近人类平均表现,但在长文本中准确率下降。错误分析揭示关键失败模式:过度依赖情绪关键词,以及在中性文本中产生误报。
原文摘要 · Abstract (English)
We introduce the SEER (Span-based Emotion Evidence Retrieval) Benchmark to test Large Language Models' (LLMs) ability to identify the specific spans of text that express emotion. Unlike traditional emotion recognition tasks that assign a single label to an entire sentence, SEER targets the underexplored task of emotion evidence detection: pinpointing which exact phrases convey emotion. This span-level approach is crucial for applications like empathetic dialogue and clinical support, which need to know how emotion is expressed, not just what the emotion is. SEER includes two tasks: identifying emotion evidence within a single sentence, and identifying evidence across a short passage of five consecutive sentences. It contains new annotations for both emotion and emotion evidence on 1200 real-world sentences. We evaluate 14 open-source LLMs and find that, while some models approach average human performance on single-sentence inputs, their accuracy degrades in longer passages. Our error analysis reveals key failure modes, including overreliance on emotion keywords and false positives in neutral text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。