arXiv:2608.22399cs.LG2026-08

分离说话人交互与情绪演化,提升对话情感识别准确率。

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

论文配图:Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
图 1 · 摘自论文原文
  • 分尺度建模:帧级与对话级双状态空间编码
  • 说话人独立动态CRF捕捉情绪演变,准确率75.81%(IEMOCAP UA)
  • 适合需要细粒度情绪分析的对话系统研究者

对话中语音情感识别需在时间尺度上融合声学线索,并处理跨说话人上下文影响与单说话人情绪演化两种交互过程。我们提出DSSM-CRF,一种纯音频架构,显式分离这两种机制。双向状态空间模型在帧级与对话级编码融合的自监督语音表示,使每个话语表征同时包含局部语调与所有说话人的上下文信息。解码器将每位说话人的话语序列构造成独立的动态条件随机场链。连续话语对的得分结合语料级转移矩阵与由两个上下文表示预测的残差。辅助目标监督每对是否情绪变化,但不参与Viterbi推理。因此,对话轮次影响上下文情感得分,却不作为另一说话人情绪轨迹的转移。DSSM-CRF在IEMOCAP上取得75.81%的未加权准确率(UA)和74.90%的加权准确率(WA),在MELD上达到54.72%的加权准确率(WA)与49.31%的加权F1(WF1)。对照实验验证了说话人分解与CRF建模的互补增益。

原文摘要 · Abstract (English)

Conversational speech emotion recognition must reconcile acoustic evidence across temporal scales with two interaction processes: cross-speaker contextual influence and within-speaker emotion evolution. We propose DSSM-CRF, an audio-only architecture that explicitly separates these processes. Bidirectional state-space models encode fused self-supervised speech representations at frame and dialogue scales, so each utterance representation captures local prosody and context from all speakers. The decoder then orders each speaker's utterances into an independent dynamic conditional random field chain. Consecutive utterances in a speaker's chain form a transition pair whose score combines a corpus-level transition matrix with a residual predicted from the two contextualized utterances. An auxiliary objective supervises whether each pair changes emotion but does not participate in Viterbi inference. Thus, interlocutor turns affect contextual emotion scores without being treated as transitions in another speaker's emotion trajectory. DSSM-CRF achieves 75.81% UA and 74.90% WA on IEMOCAP, and 54.72% WA and 49.31% WF1 on MELD. Matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.

情感识别对话系统状态空间模型动态CRF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。