arXiv:2507.08344cs.CV2025-07IJCAI被引 15

融合多模态信息,精准识别微小手势,性能领先。

MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

  • 结合肢体、视觉与深度数据,通过加权融合提升识别精度。
  • 在iMiGUE数据集上达到73.213%的准确率,优于现有方法。
  • 适合智能交互、人机协同等需要精细动作识别的场景。

本文提出MM-Gesture,由HFUT-VUT团队开发,在IJCAI 2025第三届MiGA挑战赛微手势分类赛道中排名第一,性能超越现有最先进方法。该框架专为识别细微且持续时间短的微手势(MGs)设计,融合关节、肢体、RGB视频、泰勒级数视频、光流视频及深度视频等多种模态的互补信息。采用PoseConv3D与Video Swin Transformer架构,并引入新颖的模态加权集成策略;通过在更大规模的MA-52数据集上进行迁移学习,进一步提升RGB模态表现。在iMiGUE基准上的大量实验,包括不同模态的消融研究,验证了方法的有效性,最终取得73.213%的顶1准确率。代码已公开于:https://github.com/momiji-bit/MM-Gesture。

原文摘要 · Abstract (English)

In this paper, we present MM-Gesture, the solution developed by our team HFUT-VUT, which ranked 1st in the micro-gesture classification track of the 3rd MiGA Challenge at IJCAI 2025, achieving superior performance compared to previous state-of-the-art methods. MM-Gesture is a multimodal fusion framework designed specifically for recognizing subtle and short-duration micro-gestures (MGs), integrating complementary cues from joint, limb, RGB video, Taylor-series video, optical-flow video, and depth video modalities. Utilizing PoseConv3D and Video Swin Transformer architectures with a novel modality-weighted ensemble strategy, our method further enhances RGB modality performance through transfer learning pre-trained on the larger MA-52 dataset. Extensive experiments on the iMiGUE benchmark, including ablation studies across different modalities, validate the effectiveness of our proposed approach, achieving a top-1 accuracy of 73.213%. Code is available at: https://github.com/momiji-bit/MM-Gesture.

手势识别多模态融合微手势视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。