提出自适应多尺度注意力网络,提升微表情识别精度。
AHMSA-Net: Adaptive Hierarchical Multi-Scale Attention Network for Micro-Expression Recognition
- 通过动态调整特征图尺寸,捕捉微表情的细粒度与粗粒度变化。
- 融合多尺度通道与空间特征,有效建模瞬时微表情动作信息。
- 在多个数据集上达78.21%准确率,适合细粒度动作识别研究者。
微表情识别(MER)因动作短暂且细微而极具挑战性。近年来基于注意力机制的深度学习方法取得一定进展,但仍存在特征捕获不足与动态适应性差的问题。为此,本文提出自适应分层多尺度注意力网络(AHMSA-Net)。首先利用微表情序列的起始帧与峰值帧生成三维光流图,包括水平光流、垂直光流和光流应变。随后将光流特征图输入AHMSA-Net,该网络由自适应分层框架与多尺度注意力机制组成。自适应分层框架通过动态调整每层特征图大小,从不同粒度(精细与粗糙)捕捉微表情的细微变化;多尺度注意力机制则通过融合不同尺度(通道与空间)的特征,学习微表情的动作信息。两者协同提升识别精度。大量实验表明,所提方法在主流微表情数据库上表现优异,在复合数据集(SMIC、SAMM、CASMEII)上达到78.21%准确率,在CASME^{}3数据集上达77.08%。
原文摘要 · Abstract (English)
Micro-expression recognition (MER) presents a significant challenge due to the transient and subtle nature of the motion changes involved. In recent years, deep learning methods based on attention mechanisms have made some breakthroughs in MER. However, these methods still suffer from the limitations of insufficient feature capture and poor dynamic adaptation when coping with the instantaneous subtle movement changes of micro-expressions. Therefore, in this paper, we design an Adaptive Hierarchical Multi-Scale Attention Network (AHMSA-Net) for MER. Specifically, we first utilize the onset and apex frames of the micro-expression sequence to extract three-dimensional (3D) optical flow maps, including horizontal optical flow, vertical optical flow, and optical flow strain. Subsequently, the optical flow feature maps are inputted into AHMSA-Net, which consists of two parts: an adaptive hierarchical framework and a multi-scale attention mechanism. Based on the adaptive downsampling hierarchical attention framework, AHMSA-Net captures the subtle changes of micro-expressions from different granularities (fine and coarse) by dynamically adjusting the size of the optical flow feature map at each layer. Based on the multi-scale attention mechanism, AHMSA-Net learns micro-expression action information by fusing features from different scales (channel and spatial). These two modules work together to comprehensively improve the accuracy of MER. Additionally, rigorous experiments demonstrate that the proposed method achieves competitive results on major micro-expression databases, with AHMSA-Net achieving recognition accuracy of up to 78.21% on composite databases (SMIC, SAMM, CASMEII) and 77.08% on the CASME^{}3 database.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。