arXiv:2510.15466cs.CV2025-10被引 1

通过分阶段时序增强,提升微表情识别的准确率。

Improving Micro-Expression Recognition with Phase-Aware Temporal Augmentation

  • 将微表情分为起始到峰值和峰值到结束两阶段,分别生成动态图像。
  • 在CASME-II和SAMM数据集上提升10%相对准确率,改善类不平衡问题。
  • 方法简单通用,适合数据稀缺场景下的微表情研究。

微表情是持续时间通常不足半秒的短暂、无意识面部动作,能揭示真实情绪,在心理学、安全与行为分析中至关重要。尽管深度学习推动了微表情识别(MER)发展,但标注数据稀缺限制了模型泛化能力及运动模式多样性。现有研究多依赖翻转、旋转等空间增强,忽视对运动特征更有效的时序增强。本文提出基于动态图像的相位感知时序增强方法:不将整个表情编码为单一起始到结束的动态图像(DI),而是分解为起始到峰值与峰值到结束两个运动阶段,分别生成对应DI,形成双相位DI增强策略。该方法丰富了运动多样性,引入关键的互补时序线索。在CASME-II与SAMM数据集上,结合六种深度架构(包括CNN、Vision Transformer和轻量级LEARNet)的实验表明,识别准确率、未加权F1分数与未加权平均召回率均显著提升,有效缓解类别不平衡。与空间增强结合后,相对性能最高提升10%。该方法简单、模型无关,适用于低资源环境,为鲁棒且可泛化的微表情识别提供新方向。

原文摘要 · Abstract (English)

Micro-expressions (MEs) are brief, involuntary facial movements that reveal genuine emotions, typically lasting less than half a second. Recognizing these subtle expressions is critical for applications in psychology, security, and behavioral analysis. Although deep learning has enabled significant advances in micro-expression recognition (MER), its effectiveness is limited by the scarcity of annotated ME datasets. This data limitation not only hinders generalization but also restricts the diversity of motion patterns captured during training. Existing MER studies predominantly rely on simple spatial augmentations (e.g., flipping, rotation) and overlook temporal augmentation strategies that can better exploit motion characteristics. To address this gap, this paper proposes a phase-aware temporal augmentation method based on dynamic image. Rather than encoding the entire expression as a single onset-to-offset dynamic image (DI), our approach decomposes each expression sequence into two motion phases: onset-to-apex and apex-to-offset. A separate DI is generated for each phase, forming a Dual-phase DI augmentation strategy. These phase-specific representations enrich motion diversity and introduce complementary temporal cues that are crucial for recognizing subtle facial transitions. Extensive experiments on CASME-II and SAMM datasets using six deep architectures, including CNNs, Vision Transformer, and the lightweight LEARNet, demonstrate consistent performance improvements in recognition accuracy, unweighted F1-score, and unweighted average recall, which are crucial for addressing class imbalance in MER. When combined with spatial augmentations, our method achieves up to a 10\% relative improvement. The proposed augmentation is simple, model-agnostic, and effective in low-resource settings, offering a promising direction for robust and generalizable MER.

微表情识别时序增强动态图像数据稀缺

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。