让机器通过视频行为更准识别对话情绪
BeMERC: Behavior-Aware MLLM-based Framework for Multimodal Emotion Recognition in Conversation
- 引入面部微表情、肢体语言等视频行为信息增强情感识别
- 在两个基准数据集上超越现有最佳方法,提升显著
- 适合研究多模态情感计算与人机共情系统的学者
对话中的多模态情感识别(MERC)是构建共情机器的关键任务,旨在为每句对话打上情感标签。当前基于多模态大模型(MLLM)的MERC研究主要关注文本或语音特征,忽略了视频中蕴含的行为信息。与文本和音频不同,视频包含丰富的面部微表情、肢体动作和姿态,能为模型提供情感触发信号,从而提升预测精度。本文提出一种新型行为感知的MLLM框架BeMERC,将细微的面部微表情、身体语言和姿势等行为信息融入基础的MLLM架构中,以更好地建模对话过程中的情感动态变化。此外,BeMERC采用两阶段指令微调策略,使模型适配对话场景,实现端到端的MERC预测训练。实验表明,BeMERC在两个基准数据集上均优于现有最先进方法,并深入探讨了视频行为信息在MERC中的关键作用。
原文摘要 · Abstract (English)
Multimodal emotion recognition in conversation (MERC), the task of identifying the emotion label for each utterance in a conversation, is vital for developing empathetic machines. Current MLLM-based MERC studies focus mainly on capturing the speaker's textual or vocal characteristics, but ignore the significance of video-derived behavior information. Different from text and audio inputs, learning videos with rich facial expression, body language and posture, provides emotion trigger signals to the models for more accurate emotion predictions. In this paper, we propose a novel behavior-aware MLLM-based framework (BeMERC) to incorporate speaker's behaviors, including subtle facial micro-expression, body language and posture, into a vanilla MLLM-based MERC model, thereby facilitating the modeling of emotional dynamics during a conversation. Furthermore, BeMERC adopts a two-stage instruction tuning strategy to extend the model to the conversations scenario for end-to-end training of a MERC predictor. Experiments demonstrate that BeMERC achieves superior performance than the state-of-the-art methods on two benchmark datasets, and also provides a detailed discussion on the significance of video-derived behavior information in MERC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。