arXiv:2509.00056cs.CV2025-09被引 1

提出新型视频转图像方法与注意力网络,提升微表情识别准确率。

Apex-Centered Spatio-Temporal Rank Pooling and Gradient Attention for Micro-Expression Recognition

  • 将视频序列转换为强调起始-峰值-结束阶段的MESTI图像
  • 在SMIC-HS、SAMM等数据集上达最优性能,跨数据集测试也领先
  • 适合需要高精度微表情分析的研究与安防、心理评估场景

微表情识别因动作细微且短暂而极具挑战。传统输入模态如极值帧、光流和动态图像难以充分捕捉这些短暂面部变化,导致性能受限。本文提出微表情时空图像(MESTI),一种针对微表情重构的动态排序池化方法,将视频序列转化为单张图像,突出微表情的起始-峰值-结束时序模式。同时,提出微表情梯度注意力网络(MEGANet),引入梯度注意力模块,增强对细微运动特征的提取能力。通过结合MESTI与MEGANet,构建更有效的微表情识别框架。大量实验验证了MESTI的有效性,在多种常规架构下优于现有输入模态;替换已有模型输入后亦实现一致性能提升。MEGANet在SMIC-HS、SAMM数据集上达到当前最佳结果,于CASMEII数据集表现优异,并在跨数据集评估中取得领先。二者结合始终优于对比方法,证明了MESTI作为优越输入模态与MEGANet作为先进识别网络的潜力,为各类应用中的高效微表情系统提供支持。

原文摘要 · Abstract (English)

Micro-expression recognition (MER) is a challenging task due to the subtle and fleeting nature of micro-expressions. Traditional input modalities, such as Apex Frame, Optical Flow, and Dynamic Image, often fail to adequately capture these brief facial movements, resulting in suboptimal performance. In this study, we introduce the Micro-expression Spatio-Temporal Image (MESTI), a micro-expression-specific reformulation of dynamic rank pooling that transforms a video sequence into a single image while emphasizing the onset-apex-offset temporal pattern of micro-expressions. Additionally, we present the Micro-expression Gradient Attention Network (MEGANet), which incorporates a proposed Gradient Attention block to enhance the extraction of fine-grained motion features from micro-expressions. By combining MESTI and MEGANet, we aim to establish a more effective approach to MER. Extensive experiments were conducted to evaluate the effectiveness of MESTI, comparing it with existing input modalities across regular architectures. Moreover, we demonstrate that replacing the input of previously published MER networks with MESTI leads to consistent performance improvements. The performance of MEGANet is also evaluated, showing that our proposed network achieves state-of-the-art results on the SMIC-HS, SAMM and competitive performance on CASMEII datasets, it also achieves leading performance in the reported cross-dataset evaluation settings. The combination of MESTI and MEGANet consistently outperforms the compared methods. These findings underscore the potential of MESTI as a superior input modality and MEGANet as an advanced recognition network, aiming to more effective MER systems in a variety of applications.

微表情识别时空建模注意力机制视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。