arXiv:2601.00369cs.CV2026-01

兼顾手部细节与身体动作,提升骨骼动作识别精度

BHaRNet: Reliability-Aware Body-Hand Modality Expertized Networks for Fine-grained Skeleton Action Recognition

  • 双流框架融合手部与身体骨骼信息,无需置信度标注
  • 在多个数据集上实现更高准确率,噪声环境下更稳定
  • 适合需要精细动作识别的应用场景,如医疗康复分析

基于骨骼的人体动作识别已取得显著进展,但多数方法仍以身体动作为中心,忽视对细粒度识别至关重要的细微手部动作。本文提出一种概率双流框架,统一建模可靠性与多模态融合,在骨架内部与跨模态域中实现专家化学习。该框架包含三个关键组件:(1) 无校准预处理流程,直接使用原始坐标,避免标准空间变换;(2) 概率性噪声或融合机制,无需显式置信度监督即可稳定可靠性感知双流学习;(3) 从骨架到跨模态集成,将四种骨架模态(关节、骨骼、关节运动、骨骼运动)与RGB表示耦合,统一建模结构与视觉运动线索。在多个基准(NTU RGB+D 60/120、PKU-MMD、N-UCLA)及新定义的手部中心基准上全面评估,结果表明该方法在噪声和异构条件下均具一致提升与鲁棒性。

原文摘要 · Abstract (English)

Skeleton-based human action recognition (HAR) has achieved remarkable progress with graph-based architectures. However, most existing methods remain body-centric, focusing on large-scale motions while neglecting subtle hand articulations that are crucial for fine-grained recognition. This work presents a probabilistic dual-stream framework that unifies reliability modeling and multi-modal integration, generalizing expertized learning under uncertainty across both intra-skeleton and cross-modal domains. The framework comprises three key components: (1) a calibration-free preprocessing pipeline that removes canonical-space transformations and learns directly from native coordinates; (2) a probabilistic Noisy-OR fusion that stabilizes reliability-aware dual-stream learning without requiring explicit confidence supervision; and (3) an intra- to cross-modal ensemble that couples four skeleton modalities (Joint, Bone, Joint Motion, and Bone Motion) to RGB representations, bridging structural and visual motion cues in a unified cross-modal formulation. Comprehensive evaluations across multiple benchmarks (NTU RGB+D~60/120, PKU-MMD, N-UCLA) and a newly defined hand-centric benchmark exhibit consistent improvements and robustness under noisy and heterogeneous conditions.

动作识别骨骼模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。