arXiv:2512.21064cs.CV2025-12中稿 · Machine Intelligen…被引 1

通过分解与组合提升骨架动作识别效率与效果

Multimodal Skeleton-Based Action Representation Learning via Decomposition and Composition

  • 先拆解多模态特征为单模态,再对齐真实单模态数据
  • 用单模态特征自监督指导多模态表示学习,提升性能
  • 在三个数据集上实现高效高精度,适合多模态动作识别场景

多模态人体动作理解是计算机视觉的重要问题,核心挑战在于有效利用不同模态间的互补性,同时保持模型效率。现有方法多采用简单的后期融合,计算开销大;而共享主干的早期融合虽高效,但性能不佳。为此,我们提出一种自监督的多模态骨架动作表征学习框架——分解与组合(Decomposition and Composition)。该框架通过分解策略将融合后的多模态特征精细拆解为独立的单模态特征,并与各自真实的单模态标签对齐;通过组合策略将多个单模态特征整合,作为自监督信号来增强多模态表征的学习。在NTU RGB+D 60、NTU RGB+D 120和PKU-MMD II三个数据集上的大量实验表明,该方法在计算成本与模型性能之间取得了出色平衡。

原文摘要 · Abstract (English)

Multimodal human action understanding is a significant problem in computer vision, with the central challenge being the effective utilization of the complementarity among diverse modalities while maintaining model efficiency. However, most existing methods rely on simple late fusion to enhance performance, which results in substantial computational overhead. Although early fusion with a shared backbone for all modalities is efficient, it struggles to achieve excellent performance. To address the dilemma of balancing efficiency and effectiveness, we introduce a self-supervised multimodal skeleton-based action representation learning framework, named Decomposition and Composition. The Decomposition strategy meticulously decomposes the fused multimodal features into distinct unimodal features, subsequently aligning them with their respective ground truth unimodal counterparts. On the other hand, the Composition strategy integrates multiple unimodal features, leveraging them as self-supervised guidance to enhance the learning of multimodal representations. Extensive experiments on the NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD II datasets demonstrate that the proposed method strikes an excellent balance between computational cost and model performance.

动作识别多模态学习自监督骨架建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。