用超图增强的Transformer识别微动作中的情绪,效果优于现有方法。
Hybrid-supervised Hypergraph-enhanced Transformer for Micro-gesture Based Emotion Recognition
- 设计超图增强的自注意力与多尺度时序卷积模块,建模微动作行为模式。
- 在iMiGUE和SMG数据集上准确率超现有方法,最佳结果达89.3%。
- 适合做情绪识别、行为理解及微动作分析的研究者参考。
微动作是无意识的身体动作,能反映人类情绪状态,正成为行为理解和情感计算领域的新兴研究方向。然而基于微动作的情绪建模仍不充分。本文提出一种混合监督的超图增强Transformer框架,通过重构行为模式实现情绪识别。编码器和解码器分别由超图增强的自注意力与多尺度时序卷积堆叠而成;为更好捕捉微动作的细微运动,解码器引入上采样操作,以自监督方式完成重建任务。提出超图增强的自注意力模块,动态更新骨架关节间的超边,建模局部细微运动关系。同时设计浅层情绪识别头,从编码器输出中学习情绪状态,结合监督信号进行端到端联合训练。在公开数据集iMiGUE和SMG上评估,性能优于现有方法,多项指标表现最佳。
原文摘要 · Abstract (English)
Micro-gestures are unconsciously performed body gestures that can convey the emotion states of humans and start to attract more research attention in the fields of human behavior understanding and affective computing as an emerging topic. However, the modeling of human emotion based on micro-gestures has not been explored sufficiently. In this work, we propose to recognize the emotion states based on the micro-gestures by reconstructing the behavior patterns with a hypergraph-enhanced Transformer in a hybrid-supervised framework. In the framework, hypergraph Transformer based encoder and decoder are separately designed by stacking the hypergraph-enhanced self-attention and multiscale temporal convolution modules. Especially, to better capture the subtle motion of micro-gestures, we construct a decoder with additional upsampling operations for a reconstruction task in a self-supervised learning manner. We further propose a hypergraph-enhanced self-attention module where the hyperedges between skeleton joints are gradually updated to present the relationships of body joints for modeling the subtle local motion. Lastly, for exploiting the relationship between the emotion states and local motion of micro-gestures, an emotion recognition head from the output of encoder is designed with a shallow architecture and learned in a supervised way. The end-to-end framework is jointly trained in a one-stage way by comprehensively utilizing self-reconstruction and supervision information. The proposed method is evaluated on two publicly available datasets, namely iMiGUE and SMG, and achieves the best performance under multiple metrics, which is superior to the existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。