通过双视角光流分析,提升4D微表情识别准确率
Dual-View Optical Flow for 4D Micro-Expression Recognition - A Multi-Stream Fusion Attention Approach
- 从双视角提取光流,分阶段捕捉微表情运动特征
- 在4DME数据集上达0.536宏平均F1,超基线50%以上
- 适合需要高精度情绪识别的医疗与安全场景
微表情识别对情感计算至关重要,但因面部动作极短暂、强度低以及4D网格数据维度高而困难。本文提出双视角光流方法,通过两个同步视角捕获微表情序列并计算光流表征运动。流程包括视图分离与逐帧人脸裁剪以保证空间一致性,基于双视角峰值运动强度自动检测临界帧,将序列分解为起始-临界和临界-结束两阶段,分别提取水平、垂直与幅值光流通道。这些通道输入三流微注意力网络(Triple-Stream MicroAttNet),该网络采用融合注意力模块自适应加权模态特征,并用挤压-激励块增强幅值表示。训练使用焦点损失缓解类别不平衡,结合Adam优化器与早停策略。在包含24名受试者、5类情绪的多标签4DME数据集上,于IJCAI 2025 4DMR挑战赛中取得0.536宏平均UF1,超过官方基线50%以上,位居第一。消融实验表明,融合注意力与SE模块各自贡献最高达3.6点的UF1提升。结果表明,双视角、分阶段光流结合多流融合,可实现鲁棒且可解释的4D微表情识别。
原文摘要 · Abstract (English)
Micro-expression recognition is vital for affective computing but remains challenging due to the extremely brief, low-intensity facial motions involved and the high-dimensional nature of 4D mesh data. To address these challenges, we introduce a dual-view optical flow approach that simplifies mesh processing by capturing each micro-expression sequence from two synchronized viewpoints and computing optical flow to represent motion. Our pipeline begins with view separation and sequence-wise face cropping to ensure spatial consistency, followed by automatic apex-frame detection based on peak motion intensity in both views. We decompose each sequence into onset-apex and apex-offset phases, extracting horizontal, vertical, and magnitude flow channels for each phase. These are fed into our Triple-Stream MicroAttNet, which employs a fusion attention module to adaptively weight modality-specific features and a squeeze-and-excitation block to enhance magnitude representations. Training uses focal loss to mitigate class imbalance and the Adam optimizer with early stopping. Evaluated on the multi-label 4DME dataset, comprising 24 subjects and five emotion categories, in the 4DMR IJCAI Workshop Challenge 2025, our method achieves a macro-UF1 score of 0.536, outperforming the official baseline by over 50\% and securing first place. Ablation studies confirm that both the fusion attention and SE components each contribute up to 3.6 points of UF1 gain. These results demonstrate that dual-view, phase-aware optical flow combined with multi-stream fusion yields a robust and interpretable solution for 4D micro-expression recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。