arXiv:2607.20445cs.CLcs.SD2026-07被引 1

建模说话人情绪历史,让模型更懂情绪延续与变化规律。

SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations

论文配图:SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations
图 1 · 摘自论文原文
  • 用说话人情绪历史构建先验,指导后续情绪判断。
  • 在IEMOCAP上达到新最好性能,多模态下表现更稳健。
  • 适合研究情绪演进、多模态情感识别的学者使用。

对话中人类情绪具有短暂性,但往往在多个话语间持续存在。例如,我们很少瞬间从快乐切换到愤怒,情绪通常平滑演变,且具有说话人特异性:有人情绪逐渐升级,有人则慢慢冷却。此外,情绪变化常由上下文因素(如新信息或突发事件)驱动。尽管情感识别在对话中取得进展,但现有方法仍过度依赖明显信号,未充分建模这些隐含因素。尤其在多模态场景中,当信号噪声大(如面部遮挡、俚语表达或麦克风干扰)时,模型易失效。为此,本文提出说话人条件情绪先验(SCoPE),一个轻量级模块,利用每位说话人的情绪历史显式建模其情绪先验,并结合情绪转移预测机制,动态平衡先验与多模态证据。进一步设计了感知转移的融合机制,通过精度加权的对数积分,实现多模态证据与说话人先验的贝叶斯式专家乘积融合。该机制在情绪持续时依赖历史先验,在情绪可能转移时优先采纳多模态证据。实验表明,该模型在IEMOCAP数据集的多模态设置下,优于近期最先进方法。

原文摘要 · Abstract (English)

In conversations, human emotions are transient; however, they tend to persist across multiple utterances. For example, we rarely switch instantly between contrasting emotions such as happiness and anger. Instead, emotions tend to evolve smoothly, and these patterns are often speaker-specific. Some people might escalate, while others gradually cool down over time. Furthermore, when emotions change during a conversation, they are often driven by contextual factors, such as newly received information or unexpected events. Even though progress has been made in Emotion Recognition in Conversations (ERC), most existing approaches still rely heavily on overt evidence and do not sufficiently model these non-apparent factors. Especially in multimodal settings, this makes these models fragile when the signals are noisy (e.g., occluded faces, slang expressions, or microphone noise). To address these limitations, we introduce Speaker-Conditioned Priors over Emotions (SCoPE). SCoPE is a light weight module that utilizes the emotional history of each speaker and explicitly models their priors for use in subsequent emotion classification. Second, we incorporate emotion shift prediction, a well-established concept in ERC, to guide the model in balancing the priors from SCoPE and multimodal evidence. Finally, we propose a shift-aware fusion mechanism that performs precision-weighted logit integration between multimodal evidence and the speaker prior, forming a Bayesian-inspired product-of-experts formulation. This dynamic fusion allows the model to rely on historical priors when emotions persist and to prioritize multimodal evidence when shifts are likely. Experimental results show our model achieves superior performance over recent state-of-the-art models on the IEMOCAP dataset in multimodal settings.

情感识别多模态情绪演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。