揭示对话情绪识别中语境与话语标记的因果关系,发现情绪依赖上下文但效果饱和快。
Causal Emotion Recognition in Conversation: Context Saturation and Discourse-Marker Evidence
- 通过控制实验验证语境重要性,最近10-30轮对话贡献90%性能提升。
- 悲伤话语更依赖上下文,其识别准确率提升22个百分点,因缺乏局部语用线索。
- 话语标记位置与情绪相关,左边缘标记使用率低反映悲伤者更被动参与对话。
我们针对对话情绪识别中的两个长期问题展开研究:哪些建模选择显著影响性能,以及识别结果如何对应可解释的话语层面模式。在IEMOCAP上进行系统性实验,并在MELD上进行跨数据集验证。通过10次随机种子的消融实验与多重比较校正的配对显著性检验,得出三项发现:第一,对话语境是主导因素,但性能快速饱和——90%的性能提升来自最近10-30轮对话,具体取决于标签集合;第二,在仅使用话语级信息时,分层句表示表现最优,但在加入轮次级上下文后优势消失,说明对话历史已涵盖大部分句内结构信息;第三,引入外部情感词典未提升结果,表明预训练编码器已捕获多数所需情感信号。在严格因果设置下,简单模型即达强性能(4分类82.69%;6分类加权F1 67.07%),证明无需未来轮次即可实现高精度。语言学分析发现5,286个话语标记实例中,情绪与标记位置存在显著关联(p < .0001):悲伤话语的左边缘标记使用率(21.9%)显著低于其他情绪(28–32%),符合左边缘标记与主动话语管理相关的理论。该发现与识别结果一致——悲伤最受益于对话上下文,暗示悲伤更具情境依赖性。
原文摘要 · Abstract (English)
We address two persistent gaps in Emotion Recognition in Conversation: which modeling choices materially affect performance, and how recognition findings connect to interpretable discourse-level patterns. We study both through a systematic investigation on IEMOCAP with cross-dataset validation on MELD. For recognition, we run controlled ablations with 10 random seeds and paired significance tests with multiple-comparisons correction, yielding three findings. First, conversational context is the dominant factor, but performance saturates quickly: roughly 90% of the gain is captured within the most recent 10-30 preceding turns, depending on the label set. Second, hierarchical sentence representations help most in utterance-only settings and show a clear advantage on MELD, but their benefit disappears once turn-level context is available, suggesting that conversational history subsumes much of the intra-utterance structure. Third, integrating an external affective lexicon does not improve results, consistent with pretrained encoders already capturing most of the affective signal needed for ERC. Under a strictly causal setting, our simple models achieve strong performance (82.69% 4-way; 67.07% 6-way weighted F1), showing that competitive accuracy is achievable without future turns. For linguistic analysis, we examine 5,286 discourse-marker occurrences and find a reliable association between emotion and marker position (p < .0001). Sad utterances show reduced left-periphery marker usage (21.9%) relative to other emotions (28-32%), consistent with accounts linking left-periphery markers to active discourse management. This aligns with our recognition results, where Sad benefits most from conversational context (+22 percentage points), suggesting sadness may be more context-dependent than emotions with stronger local pragmatic cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。