通过记忆机制建模情绪共现模式,提升多情绪识别准确率。
Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition
- 构建原型记忆库融合多模态信号,捕捉跨模态语义关联。
- 在两个公开数据集上优于当前最佳方法,显著提升混合情绪预测精度。
- 适合关注情感计算、多情绪识别的研究者与应用开发者。
从多模态生理与行为信号中进行情绪识别在情感计算中至关重要,但现有模型大多局限于实验室环境下单一情绪的预测。真实世界中情绪常同时存在多种状态,促使研究转向将混合情绪识别视为情绪分布学习问题。然而,当前方法往往忽略共现情绪间的效价一致性与结构化关联。为此,我们提出记忆引导的原型共现学习(MPCL)框架,显式建模情绪共现模式。首先,通过多尺度关联记忆机制融合多模态信号;为捕捉跨模态语义关系,构建特定情绪的原型记忆库,生成丰富的生理与行为表征,并采用原型关系蒸馏确保潜在原型空间中的跨模态对齐。此外,受人类认知记忆系统启发,引入记忆检索策略,提取情绪类别间的语义级共现关联。通过自下而上的层级抽象过程,模型学习到具有情感信息的表征,实现精准的情绪分布预测。在两个公开数据集上的综合实验表明,MPCL在混合情绪识别上持续优于现有最先进方法,且在定量与定性层面均表现优异。
原文摘要 · Abstract (English)
Emotion recognition from multi-modal physiological and behavioral signals plays a pivotal role in affective computing, yet most existing models remain constrained to the prediction of singular emotions in controlled laboratory settings. Real-world human emotional experiences, by contrast, are often characterized by the simultaneous presence of multiple affective states, spurring recent interest in mixed emotion recognition as an emotion distribution learning problem. Current approaches, however, often neglect the valence consistency and structured correlations inherent among coexisting emotions. To address this limitation, we propose a Memory-guided Prototypical Co-occurrence Learning (MPCL) framework that explicitly models emotion co-occurrence patterns. Specifically, we first fuse multi-modal signals via a multi-scale associative memory mechanism. To capture cross-modal semantic relationships, we construct emotion-specific prototype memory banks, yielding rich physiological and behavioral representations, and employ prototype relation distillation to ensure cross-modal alignment in the latent prototype space. Furthermore, inspired by human cognitive memory systems, we introduce a memory retrieval strategy to extract semantic-level co-occurrence associations across emotion categories. Through this bottom-up hierarchical abstraction process, our model learns affectively informative representations for accurate emotion distribution prediction. Comprehensive experiments on two public datasets demonstrate that MPCL consistently outperforms state-of-the-art methods in mixed emotion recognition, both quantitatively and qualitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。