arXiv:2609.03394cs.CL2026-09

构建对比情绪数据集,让模型理解同一事件引发相反情绪。

Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

论文配图:Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory
图 1 · 摘自论文原文
  • 基于评估理论设计1000条人类标注语句,每句含对立情绪触发场景。
  • 最强大模型仅达67.3%宏平均F1,远低于人类一致水平。
  • 数据集可作训练信号,提升下游情绪识别性能。

情感识别基准通常仅预测单个情绪,忽略了同一事件引发不同情绪的现实场景。例如,孩子因兴奋踢前排座椅,而前方乘客则感到愤怒。我们提出CHIARO,一个基于评估理论的1000条人工标注句子基准,每个场景描述单一因果触发事件,引发一人正向情绪、另一人负向情绪,涵盖十类情绪分类。我们对七种前沿大模型和四种现成情绪分类器进行评测,最强模型仅达67.3%宏平均F1,远低于人类一致性水平,现有分类器得分接近随机。除评估外,CHIARO还可作为训练信号:结合已有情绪语料后,下游分类器在自身及十个外部基准上表现均优于原始模型,证明其为情感识别的互补训练资源。

原文摘要 · Abstract (English)

Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.

情绪识别对比学习评估理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。