LLM情绪标注只学到了明确词汇对应的情绪,没学会人类判断的不确定性。
LLMs Capture Emotion Labels, Not Emotion Uncertainty: Distributional Analysis and Calibration of Human-LLM Judgment Gaps

- 通过对比人类与LLM在情绪分布上的差异,发现零样本模型偏离严重
- 微调后模型才接近人类分布,且关键在领域适配而非模型大小
- 提出三种轻量校准方法,可降低14%分布差距,指导何时可用LLM替代人工
人类标注者对情绪标签常有分歧,但多数大语言模型(LLM)情绪评估将判断简化为单一标准答案,丢失了分歧所蕴含的分布信息。本文通过比较四个零样本LLM及一个微调后的RoBERTa基线,在GoEmotions与EmoBank两个互补数据集上共64万条响应中的人类与LLM情绪判断分布,发现零样本模型与人类分布显著偏离;而领域内微调而非模型规模,是缩小差距的关键。我们提出一种基于词汇锚定的量化透明度评分,可预测各情绪类别的人类-LLM一致性:LLM能可靠捕捉具显性词汇标记的情绪,但在依赖语境推断的复杂情绪上系统性失败,该模式在分类与连续情绪框架中均成立。进一步提出三种轻量级后处理校准方法,使分布差距减少最高达14%,并提供实用指南,明确指出在何种情况下可使用LLM标注替代人工。
原文摘要 · Abstract (English)
Human annotators frequently disagree on emotion labels, yet most evaluations of Large Language Model (LLM) emotion annotation collapse these judgments into a single gold standard, discarding the distributional information that disagreement encodes. We ask whether LLMs capture the structure of this disagreement, not just majority labels, by comparing emotion judgment distributions between human annotators and four zero-shot LLMs, plus a fine-tuned RoBERTa baseline, across two complementary benchmarks: GoEmotions and EmoBank, totaling 640,000 LLM responses. Zero-shot models diverge substantially from human distributions, and in-domain fine-tuning, not model scale, is required to close the gap. We formalize a lexical-grounding gradient through a quantitative transparency score that predicts per-category human--LLM agreement: LLMs reliably capture emotions with explicit lexical markers but systematically fail on pragmatically complex emotions requiring contextual inference, a pattern that replicates across both categorical and continuous emotion frameworks. We further propose three lightweight post-hoc calibration methods that reduce the distributional gap by up to 14\%, and provide actionable guidelines for when LLM emotion annotations can, and cannot, substitute for human labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。