arXiv:2606.31249cs.CV2026-06被引 2

提出多模态时序建模框架,提升长视频隐性情绪识别准确率至76.9%

Rethinking the Role of Feature Engineering and Learning Strategies in Few-Shot Hidden Emotion Recognition

论文配图:Rethinking the Role of Feature Engineering and Learning Strategies in Few-Shot Hidden Emotion Recognition
图 1 · 摘自论文原文
  • 用跨注意力机制融合静态姿态与动态微动作特征,消除身份偏差
  • 在长视频中实现76.9%测试准确率,优于前代方案
  • 揭示视觉大模型在微动态任务中的表征坍塌问题,适合情绪识别研究者

本文介绍我们团队XInsight Lab在第四届EI-MIGA-IJCAI挑战赛第三赛道中获得第一名的解决方案,测试准确率达0.76923。针对长视频中隐性情绪证据弱且稀疏的问题,本文扩展了上届竞赛的获胜方案,提出一种紧凑的多模态时序建模框架。该框架整合并评估了多种源特征:2D/3D骨骼、面部表情Blendshapes、DINOv2/v3视觉基础模型、X-CLIP视频特征及Gemini语义先验。架构上,提出一种交叉注意力机制,以静态姿态特征(Base)为查询,动态微运动差分特征(Offset)为键和值,通过捕捉局部相对速度,消除个体体型与身份相关的静态偏差。同时,采用基于多实例学习的自适应池化方法,在长序列中提取瞬时情绪并抑制背景噪声。最后,论文揭示了通用视觉基础模型在微动态任务中存在表征坍塌现象,并分析其机制:网络因捷径学习与死记硬背,陷入由公共排行榜驱动的伪泛化。

原文摘要 · Abstract (English)

In this paper, we present the solution developed by our team, XInsight Lab, which achieved first place in Track 3 of the 4th EI-MIGA-IJCAI Challenge with a test accuracy of 0.76923. To address the challenge of weak and sparse implicit emotion evidence in long videos, this paper extends the winning solution from the previous competition and proposes a compact multi-modal temporal modeling framework. The framework integrates and evaluates the effects of multi-source features, including 2D/3D skeletons, facial expression Blendshapes, DINOv2/v3 vision foundation models, X-CLIP video features, and Gemini semantic priors. Architecturally, we propose a cross-attention mechanism that utilizes static pose features, denoted as Base, as the Query and dynamic micro-motion differential features, denoted as Offset, as the Key and Value. By capturing local relative velocities, this mechanism eliminates static biases related to individual body shape and identity. Concurrently, an adaptive pooling method based on Multiple Instance Learning is employed to extract instantaneous emotions while suppressing background noise in long sequences. Finally, the paper reveals the representation collapse phenomenon of general vision foundation models in micro-dynamic tasks, and analyzes the underlying mechanisms where networks fall into public-leaderboard-driven pseudo-generalization due to shortcut learning and rote memorization.

情绪识别多模态小样本微动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。