arXiv:2409.11906cs.RO2024-09被引 5

融合热成像、表情动作与文本上下文,提升情绪识别准确率

Fusion in Context: A Multimodal Approach to Affective State Recognition

  • 用Transformer融合热成像、表情动作和文本三模态数据
  • 在桌游实验中实现比单模态更高的情绪识别精度
  • 适合人机交互与情感计算研究者参考

准确识别人类情绪是情感计算与人机交互中的关键挑战。情绪状态对行为、决策和社会互动具有重要影响,但其表达受上下文因素干扰,忽略上下文易导致误判。多模态融合(如面部表情、语音和生理信号)在提升情绪识别方面展现出潜力。本文提出一种基于Transformer的多模态融合方法,结合面部热成像数据、面部动作单元(Facial Action Units)和文本上下文信息,实现上下文感知的情绪识别。采用模态专用编码器学习各模态特有表征,通过加性融合后由共享Transformer编码器捕捉时序依赖与模态间交互。在参与者参与设计用于诱发多种情绪状态的实体桌面版吃豆人游戏所收集的数据集上进行评估。结果表明,引入上下文信息与多模态融合能显著提升情绪识别效果。

原文摘要 · Abstract (English)

Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional expressions can be influenced by contextual factors, leading to misinterpretations if context is not considered. Multimodal fusion, combining modalities like facial expressions, speech, and physiological signals, has shown promise in improving affect recognition. This paper proposes a transformer-based multimodal fusion approach that leverages facial thermal data, facial action units, and textual context information for context-aware emotion recognition. We explore modality-specific encoders to learn tailored representations, which are then fused using additive fusion and processed by a shared transformer encoder to capture temporal dependencies and interactions. The proposed method is evaluated on a dataset collected from participants engaged in a tangible tabletop Pacman game designed to induce various affective states. Our results demonstrate the effectiveness of incorporating contextual information and multimodal fusion for affective state recognition.

情绪识别多模态融合人脸热成像上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。