让大模型精准识别细微手势,关键在动态聚焦运动区域。
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

- 用门控机制动态提取动作相关区域和帧间运动差异
- 在iMiGUE和SMG数据集上分别达67.32%和73.11%准确率
- 适合需要高精度微手势分析的医疗、人机交互场景
微手势识别需捕捉短暂且空间局部化的运动,常被静态外观和背景噪声掩盖。尽管多模态大语言模型(MLLMs)在通用视频理解上表现优异,但其对细微运动感知能力不足,且依赖静态姿态先验。为此,我们提出GMoT——一种门控运动感知分词模块,将稀疏的运动证据提炼为紧凑序列,用于后续时序建模。GMoT通过空间加权池化动态聚焦动作相关区域,利用相邻帧差分捕捉精确运动能量,并采用保守初始化的语义门自适应融合这些线索至视觉流。为进一步实现基于证据的推理,我们引入渐进式奖励引导策略优化框架,结合半监督标注流程生成解剖学聚焦描述。在iMiGUE(67.32%)和SMG(73.11%)数据集上均优于对比方法,分别提升Qwen3-VL-8B基线6.80和3.11个百分点。我们还引入身体区域定位(BRG)召回率作为解剖定位代理指标,并设计跨域迁移协议。实验表明,增强模型在域内准确率、抗标签扰动及小样本跨域迁移方面均有提升,且生成推理过程保持高解剖一致性。
原文摘要 · Abstract (English)
Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. While Multimodal Large Language Models (MLLMs) excel at general video understanding, they inherently struggle with subtle kinematics and often rely on static posture priors. To this end, we propose GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling. GMoT dynamically spotlights action-relevant regions via spatially weighted pooling, extracts adjacent-frame temporal differencing to capture precise motion energy, and adaptively fuses these cues into the visual stream using a conservatively initialized semantic gate. To transition from simple classification to evidence-grounded reasoning, we further introduce a progressive reward-guided policy refinement paradigm, supported by a semi-supervised annotation pipeline that generates anatomically focused captions. Beyond achieving the best Top-1 accuracy among the compared methods on iMiGUE (67.32\%) and SMG (73.11\%), improving the Qwen3-VL-8B baseline by +6.80 and +3.11 points, our framework introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions, together with an overlapping-label cross-domain transfer protocol between iMiGUE and SMG. Extensive evaluations demonstrate that our GMoT-augmented model improves in-domain accuracy, retains clear gains under label-preserving corruptions, and improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding in its generated rationales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。