通过自监督放大微表情运动并稀疏建模关键区域,提升识别精度。
AMMSM: Adaptive Motion Magnification and Sparse Mamba for Micro-Expression Recognition
- 自监督放大微表情运动信号,增强细微变化捕捉能力。
- 稀疏空间选择Mamba架构在两个数据集上达最新最佳准确率。
- 适合关注情绪识别与视频分析的算法研究者。
微表情是人真实情绪的无意识表现,但持续时间短、信号微弱,给下游识别带来挑战。本文提出多任务学习框架AMMSM,通过自监督方式增强微表情的细微运动信号,同时采用稀疏空间选择的Mamba架构,结合先进的视觉Mamba模型,更有效地建模关键运动区域及其特征表示。此外,利用进化搜索优化运动放大系数和空间稀疏比例,再通过微调进一步提升性能。在两个标准数据集上的大量实验表明,所提方法在准确率和鲁棒性方面均达到当前最优水平。
原文摘要 · Abstract (English)
Micro-expressions are typically regarded as unconscious manifestations of a person's genuine emotions. However, their short duration and subtle signals pose significant challenges for downstream recognition. We propose a multi-task learning framework named the Adaptive Motion Magnification and Sparse Mamba (AMMSM) to address this. This framework aims to enhance the accurate capture of micro-expressions through self-supervised subtle motion magnification, while the sparse spatial selection Mamba architecture combines sparse activation with the advanced Visual Mamba model to model key motion regions and their valuable representations more effectively. Additionally, we employ evolutionary search to optimize the magnification factor and the sparsity ratios of spatial selection, followed by fine-tuning to improve performance further. Extensive experiments on two standard datasets demonstrate that the proposed AMMSM achieves state-of-the-art (SOTA) accuracy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。