通过增强数据与时空注意力,精准识别视频中的微手势位置与类别。
Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention
- 手工设计数据增强提升模型泛化能力
- 结合时空注意力机制实现高精度定位与分类
- 在挑战赛中取得38.03的F1分数,排名第一
本文介绍我们团队为IJCAI 2025 MiGA挑战赛中微手势在线识别任务提出的最新解决方案HFUT-VUT。该任务旨在从未剪辑视频中定位多个微手势实例的时间位置并识别其类别,相比传统动作检测更强调细粒度分类与精确起止时间判定。由于微手势是自发性人类行为,类别间差异显著,识别难度大。为此,我们提出手工数据增强与时空注意力机制,以提升模型对微手势的分类与定位能力。最终方案在测试集上获得38.03的F1分数,较此前最优方法提升37.9%,并在该赛道中位列第一。
原文摘要 · Abstract (English)
In this paper, we introduce the latest solution developed by our team, HFUT-VUT, for the Micro-gesture Online Recognition track of the IJCAI 2025 MiGA Challenge. The Micro-gesture Online Recognition task is a highly challenging problem that aims to locate the temporal positions and recognize the categories of multiple micro-gesture instances in untrimmed videos. Compared to traditional temporal action detection, this task places greater emphasis on distinguishing between micro-gesture categories and precisely identifying the start and end times of each instance. Moreover, micro-gestures are typically spontaneous human actions, with greater differences than those found in other human actions. To address these challenges, we propose hand-crafted data augmentation and spatial-temporal attention to enhance the model's ability to classify and localize micro-gestures more accurately. Our solution achieved an F1 score of 38.03, outperforming the previous state-of-the-art by 37.9%. As a result, our method ranked first in the Micro-gesture Online Recognition track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。