arXiv:2508.13483cs.CV2025-08中稿 · IJCNN 2025被引 1

融合2D与3D特征,用多任务学习提升微表情识别准确率

FAMNet: Integrating 2D and 3D Features for Micro-expression Recognition via Multi-task Learning and Hierarchical Attention

  • 通过双流结构整合2D和3D CNN,提取时空联合特征
  • 在SAMM、CASME II等数据集上达到83.75% (UAR)的识别率
  • 多任务学习促进微表情与面部动作单元协同优化

微表情识别(MER)在多个领域具有重要应用价值,但其持续时间短、强度低,给识别带来巨大挑战。现有深度学习方法主要采用静态图像、动态图像序列或双流融合方式加载数据,难以有效提取微表情的细粒度时空特征。本文提出一种基于多任务学习与分层注意力机制的MER新方法——FAMNet,通过融合2D CNN(AMNet2D)与3D CNN(AMNet3D)构建双流网络,两者共享ResNet18主干网络与注意力模块。训练时分别采用不同数据加载策略适配两路网络,联合进行微表情识别与面部动作单元检测(FAUD)任务,并采用参数硬共享实现信息关联,显著提升识别效果。实验表明,FAMNet在SAMM、CASME II和MMEW数据集上分别取得83.75%(UAR)和84.03%(UF1)的性能;在更具挑战性的CAS(ME)³数据集上也达到51%(UAR)和43.42%(UF1)。

原文摘要 · Abstract (English)

Micro-expressions recognition (MER) has essential application value in many fields, but the short duration and low intensity of micro-expressions (MEs) bring considerable challenges to MER. The current MER methods in deep learning mainly include three data loading methods: static images, dynamic image sequence, and a combination of the two streams. How to effectively extract MEs' fine-grained and spatiotemporal features has been difficult to solve. This paper proposes a new MER method based on multi-task learning and hierarchical attention, which fully extracts MEs' omni-directional features by merging 2D and 3D CNNs. The fusion model consists of a 2D CNN AMNet2D and a 3D CNN AMNet3D, with similar structures consisting of a shared backbone network Resnet18 and attention modules. During training, the model adopts different data loading methods to adapt to two specific networks respectively, jointly trains on the tasks of MER and facial action unit detection (FAUD), and adopts the parameter hard sharing for information association, which further improves the effect of the MER task, and the final fused model is called FAMNet. Extensive experimental results show that our proposed FAMNet significantly improves task performance. On the SAMM, CASME II and MMEW datasets, FAMNet achieves 83.75% (UAR) and 84.03% (UF1). Furthermore, on the challenging CAS(ME)$^3$ dataset, FAMNet achieves 51% (UAR) and 43.42% (UF1).

微表情识别多任务学习时空特征注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。