arXiv:2604.00819cs.CLcs.AI2026-04

构建多维情感数据集,用贝叶斯推理提升情感预测一致性。

Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding

  • 基于普拉奇克情绪理论构建8维情感向量标注的场景数据集
  • 轻量级贝叶斯框架使情感共现统计融入联合推断,准确率提升2.24%
  • 适合研究情感结构依赖与大模型多维情绪理解能力的学者

自然语言中的情绪理解本质上是多维度推理问题,多个情感信号通过语境、人际关联和情境线索相互作用。然而,现有情绪理解基准大多依赖短文本和预定义标签,将过程简化为独立标签预测,忽略了情感间的结构性依赖。为此,我们提出了情感场景(EmoScene),一个基于理论的基准数据集,包含4,731个上下文丰富的场景,标注了源自普拉奇克基本情绪的8维情感向量。受情感极少独立出现的启发,我们进一步提出一种考虑情感共现统计的纠缠感知贝叶斯推断框架,实现对情感向量的联合后验推断。该轻量级后处理无需参数更新,提升了预测的结构一致性,在不增加任何成本的情况下,整体词汇准确率提升2.24%。EmoScene为研究多维情绪理解及当前语言模型的局限性提供了具有挑战性的基准。

原文摘要 · Abstract (English)

Understanding emotions in natural language is inherently a multi-dimensional reasoning problem, where multiple affective signals interact through context, interpersonal relations, and situational cues. However, most existing emotion understanding benchmarks rely on short texts and predefined emotion labels, reducing this process to independent label prediction and ignoring the structured dependencies among emotions. To address this limitation, we introduce Emotional Scenarios (EmoScene), a theory-grounded benchmark of 4,731 contextrich scenarios annotated with an 8-dimensional emotion vector derived from Plutchik's basic emotions. Motivated by the observation that emotions rarely occur independently, we further propose an entanglement-aware Bayesian inference framework that incorporates emotion co-occurrence statistics to perform joint posterior inference over the emotion vector. This lightweight post-processing does not require any parameter updates and improves the structural consistency of predictions, and yields overall gains of 2.24% Lexical Accuracy without any additional cost. EmoScene therefore provides a challenging benchmark for studying multi-dimensional emotion understanding and the limitations of current language models.

多维情绪贝叶斯推理情感建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。