arXiv:2605.12838cs.AI2026-05

用多模态数据追踪对话中的持续情绪变化,更轻量且可解释。

Multimodal Hidden Markov Models for Persistent Emotional State Tracking

论文配图:Multimodal Hidden Markov Models for Persistent Emotional State Tracking
图 1 · 摘自论文原文
  • 基于多模态情感特征构建隐马尔可夫模型,捕捉对话中持续的情绪阶段。
  • 在临床数据上验证,能可靠识别有意义的情绪周期并提升大模型响应质量。
  • 计算成本远低于大模型方法,适合大规模实时情绪分析应用。

通过整体分析对话中单个语句的情感,实现对对话情绪演变过程的可解释追踪,是理解与引导实际沟通(尤其是临床场景)的核心。现有情绪识别方法仅关注单句层面,忽略了真实对话中持续存在的情绪阶段。本文提出一种轻量级框架,利用粘性因子化HDP-HMM对视频、音频与文本同步提取的多模态效价-唤醒度表示进行建模,将对话情绪视为一系列潜在情绪状态序列。通过大语言模型评分、几何与时间一致性指标评估,结果表明该模型生成的阶段序列比基线高斯HMM更具可解释性,且计算成本仅为基于大语言模型的对话状态追踪方法的一小部分。在临床数据集上的问答实验进一步显示,从多模态效价-唤醒度轨迹中可有效恢复有意义的情绪阶段,并通过上下文增强显著改善大模型在情绪不稳状态下的输出质量。该框架为大规模可解释、轻量化、可操作的情绪动态分析提供了可行路径。

原文摘要 · Abstract (English)

Tracking an interpretable emotional arc of a conversation via the sentiment of individual utterances processed as a whole is central to both understanding and guiding communication in applied, especially clinical, conversational contexts. Existing approaches to emotion recognition operate at the utterance level, obscuring the persistent phases that characterize real conversational dynamics. We propose a lightweight framework that models conversational emotion as a sequence of latent emotional regimes using sticky factorial HDP-HMMs over multimodal valence-arousal representations derived from simultaneous video, audio and textual input. We evaluate the quality of regime prediction using LLM-as-a-Judge, geometric, and temporal consistency metrics, demonstrating that the sticky HDP-HMM produces more interpretable regime sequences than the baseline Gaussian HMM at a fraction of the computational cost of LLM-based dialogue state tracking methods. In addition, Question-Answer experiments in a clinical dataset suggest that meaningful emotional phases can reliably be recovered from multimodal valence-arousal trajectories and used to improve the quality of LLM responses in unstable affective regimes via context augmentation. This framework thus opens a path toward interpretable, lightweight, and actionable analysis of conversational emotion dynamics at scale.

情绪追踪多模态隐马尔可夫临床对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。