针对微动作识别难题,提出分部位专家模型,提升细微动作辨识能力。
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
- 分区域专家分工处理头、躯干、上下肢运动信息
- 在三个数据集上实现最优性能,尤其改善模糊与低幅动作识别
- 适合关注人体细粒度动作分析的研究者
微动作(如瞥视、点头、轻微姿态变化)持续时间短、幅度小,蕴含丰富社交意义,但现有动作识别模型难以捕捉。本文提出B-MoE,一种身体部位感知的混合专家框架,显式建模人体运动结构。每个专家专注特定身体区域(头、躯干、上肢、下肢),基于轻量级宏-微运动编码器(M3E)捕捉长程上下文与细粒度局部运动。通过交叉注意力路由机制,动态学习区域间关系并选择最相关信息。双流编码器融合区域语义与全局运动特征,联合捕捉空间局部线索与时间细微变化。在MA-52、SocialGesture和MPII-GroupInteraction三个挑战性数据集上均取得显著领先,尤其在模糊、低频及低幅类别上表现突出。
原文摘要 · Abstract (English)
Micro-actions, fleeting and low-amplitude motions, such as glances, nods, or minor posture shifts, carry rich social meaning but remain difficult for current action recognition models to recognize due to their subtlety, short duration, and high inter-class ambiguity. In this paper, we introduce B-MoE, a Body-part-aware Mixture-of-Experts framework designed to explicitly model the structured nature of human motion. In B-MoE, each expert specializes in a distinct body region (head, body, upper limbs, lower limbs), and is based on the lightweight Macro-Micro Motion Encoder (M3E) that captures long-range contextual structure and fine-grained local motion. A cross-attention routing mechanism learns inter-region relationships and dynamically selects the most informative regions for each micro-action. B-MoE uses a dual-stream encoder that fuses these region-specific semantic cues with global motion features to jointly capture spatially localized cues and temporally subtle variations that characterize micro-actions. Experiments on three challenging benchmarks (MA-52, SocialGesture, and MPII-GroupInteraction) show consistent state-of-theart gains, with improvements in ambiguous, underrepresented, and low amplitude classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。