arXiv:2512.23291cs.CV2025-12IJCAI被引 1

通过多模态融合提升微表情与情绪识别精度

Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition

  • 融合视频与骨骼数据,用跨模态令牌融合捕捉细微动作
  • 在iMiGUE数据集上实现情绪识别第二名,准确率显著提升
  • 适合做行为分析、人机交互与情感计算的研究者参考

微表情识别与基于行为的情绪预测都是极具挑战的任务,需建模细微的细粒度人类行为,主要依赖视频和骨骼姿态数据。本文提出两种多模态框架,在iMiGUE数据集上同时解决这两个问题。针对微表情分类,探索了RGB与3D姿态表征的互补优势,利用MViTv2-S提取视频嵌入,2s-AGCN提取骨骼嵌入,并通过跨模态令牌融合模块整合空间与姿态信息。对于情绪识别,扩展为基于行为的情绪预测任务(二分类),采用SwinFace和MViTv2-S分别提取面部与上下文嵌入,通过设计的InterFusion模块捕捉表情与肢体动作。在MiGA 2025挑战赛的iMiGUE数据集上,所提方法在行为基情绪预测任务中表现稳健,取得第二名成绩。

原文摘要 · Abstract (English)

Micro-gesture recognition and behavior-based emotion prediction are both highly challenging tasks that require modeling subtle, fine-grained human behaviors, primarily leveraging video and skeletal pose data. In this work, we present two multimodal frameworks designed to tackle both problems on the iMiGUE dataset. For micro-gesture classification, we explore the complementary strengths of RGB and 3D pose-based representations to capture nuanced spatio-temporal patterns. To comprehensively represent gestures, video, and skeletal embeddings are extracted using MViTv2-S and 2s-AGCN, respectively. Then, they are integrated through a Cross-Modal Token Fusion module to combine spatial and pose information. For emotion recognition, our framework extends to behavior-based emotion prediction, a binary classification task identifying emotional states based on visual cues. We leverage facial and contextual embeddings extracted using SwinFace and MViTv2-S models and fuse them through an InterFusion module designed to capture emotional expressions and body gestures. Experiments conducted on the iMiGUE dataset, within the scope of the MiGA 2025 Challenge, demonstrate the robust performance and accuracy of our method in the behavior-based emotion prediction task, where our approach secured 2nd place.

微表情情绪识别多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。