arXiv:2503.14960cs.CV2025-03中稿 · IROS 2025, project…被引 2

专精手部动作识别,提升细粒度行为理解精度

Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition

  • 双专家网络分离处理身体与手部动作特征
  • 跨注意力机制提升手部细微动作识别率至93.0%
  • 适合需要精准手势识别的机器人交互场景

基于骨骼的人体动作识别在机器人与人机交互中至关重要。然而,现有方法多聚焦整体动作,忽视对细粒度动作区分至关重要的细微手部运动。近期工作采用统一图结构融合身体、手部和足部关键点以捕捉动态细节,但常因身体与手部动作特性差异及空间池化导致手部细节模糊。本文提出BHaRNet(Body-Hand Action Recognition Network),在典型身体专家模型基础上增加手部专家模型,通过集成损失联合训练两路分支,实现协作专业化,类似混合专家(MoE)机制。进一步引入跨注意力机制,通过专家分支与池化注意力模块实现特征级交互,选择性融合互补信息。受MMNet启发,还拓展至多模态任务,利用RGB信息时由身体特征引导学习,捕获更丰富上下文线索。在大规模基准数据集(NTU RGB+D 60、NTU RGB+D 120、PKU-MMD、Northwestern-UCLA)上的实验表明,BHaRNet达到当前最优性能——手部密集动作识别准确率从86.4%提升至93.0%,同时参数量和计算量均低于相关统一方法。

原文摘要 · Abstract (English)

Skeleton-based Human Action Recognition (HAR) is a vital technology in robotics and human-robot interaction. However, most existing methods concentrate primarily on full-body movements and often overlook subtle hand motions that are critical for distinguishing fine-grained actions. Recent work leverages a unified graph representation that combines body, hand, and foot keypoints to capture detailed body dynamics. Yet, these models often blur fine hand details due to the disparity between body and hand action characteristics and the loss of subtle features during the spatial-pooling. In this paper, we propose BHaRNet (Body-Hand action Recognition Network), a novel framework that augments a typical body-expert model with a hand-expert model. Our model jointly trains both streams with an ensemble loss that fosters cooperative specialization, functioning in a manner reminiscent of a Mixture-of-Experts (MoE). Moreover, cross-attention is employed via an expertized branch method and a pooling-attention module to enable feature-level interactions and selectively fuse complementary information. Inspired by MMNet, we also demonstrate the applicability of our approach to multi-modal tasks by leveraging RGB information, where body features guide RGB learning to capture richer contextual cues. Experiments on large-scale benchmarks (NTU RGB+D 60, NTU RGB+D 120, PKU-MMD, and Northwestern-UCLA) demonstrate that BHaRNet achieves SOTA accuracies -- improving from 86.4\% to 93.0\% in hand-intensive actions -- while maintaining fewer GFLOPs and parameters than the relevant unified methods.

动作识别骨骼分析跨注意力细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。