通过动态连接面部区域,联合建模微表情的时空特征,提升识别准确率。
STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition

- 用光流和注意力选关键帧,结合图网络与Transformer同步捕捉时空动态。
- 在6个数据集上实现更高精度与跨数据集泛化能力,比现有方法提升显著。
- 适合关注可解释性与跨数据集鲁棒性的微表情研究者使用。
微表情识别因面部肌肉运动细微且持续时间短而具有挑战性。现有方法过度依赖峰值起始帧,忽视帧间精细动态,且分别建模空间与时间信息,限制了跨数据集泛化能力。为此,我们提出STAG,一种动态感兴趣区域-动作单元(ROI-AU)耦合的时空网络,联合建模运动流与自适应面部连接关系。该框架基于幅度选择提取判别帧的光流,并引入时间注意力机制。双分支结构融合增强图注意力网络(用于结构化空间推理)与Transformer编码器(用于时序建模)。双向交叉注意力模块实现时空特征相互优化,而AU引导的动态连接根据肌肉激活模式自适应调整面部区域间交互。Transformer捕捉超越峰值基准的细微时序动态,提升语义一致性与可解释性。融合表征采用焦点损失优化,在CASME II、4DME、DFME、NaME、SAMM和SMIC-HS六个数据集上评估。大量实验表明,该方法在鲁棒性、泛化性、可解释性及计算效率方面均优于现有方法,验证了自适应关系推理、AU引导动态连接与深层时空特征融合在跨数据集微表情识别中的有效性。
原文摘要 · Abstract (English)
Micro-expression recognition is challenging due to subtle and short-lived facial muscle movements. Existing methods rely heavily on apex-onset frames, overlook fine-grained inter-frame dynamics, and separately model spatial and temporal information, limiting generalization across datasets. To address these challenges, we propose STAG, a dynamic ROI-AU-coupled spatial-temporal network that jointly models motion flow and adaptive facial connectivity. The framework extracts optical flow from discriminative frames using magnitude-based selection and temporal attention. A dual-branch architecture combines an enhanced graph attention network for structured spatial reasoning with a transformer encoder for temporal modeling. A bidirectional cross-attention module enables mutual refinement of spatial and temporal features, while AU-guided dynamic connectivity adapts facial region interactions according to muscle activation patterns. The transformer captures subtle temporal dynamics beyond apex-based approaches, improving semantic consistency and interpretability for explainable micro-expression recognition. The fused representation is optimized using focal loss and evaluated on CASME II, 4DME, DFME, NaME, SAMM, and SMIC-HS. Extensive experiments demonstrate improved robustness, generalization, interpretability, and computational efficiency, confirming the effectiveness of adaptive relational reasoning, AU-guided dynamic connectivity, and deep spatial-temporal feature fusion for accurate cross-dataset micro-expression recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。