arXiv:2512.21916cs.CV2025-12

将人体关节区域的图像块建模为图,实现更精准的多模态动作识别。

Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition

  • 用包含人体关节的图像块构建时空图,聚焦人体中心信息。
  • 在三个数据集上达到当前最优性能,统一与分离建模均领先。
  • 支持低质量骨骼数据,适合需融合视觉与骨骼信息的场景。

尽管动作识别已取得显著进展,但融合RGB与骨骼模态的多模态方法仍受二者固有异质性制约,难以充分挖掘其互补潜力。本文提出PAN,首个面向多模态动作识别的人体中心图表示学习框架,将包含人体关节的RGB图像块嵌入表示为时空图。该人体中心建模范式抑制了RGB帧中的冗余信息,与基于骨骼的方法高度对齐,从而实现更有效且语义一致的多模态特征融合。由于令牌嵌入采样严重依赖二维骨骼数据,我们进一步提出基于注意力的后校准机制,在几乎不损失模型性能的前提下降低对高质量骨骼数据的依赖。为探索PAN与骨骼基方法的集成潜力,我们设计两种变体:PAN-Ensemble采用双路径图卷积网络结合晚期融合;PAN-Unified则在单个网络内完成统一图表示学习。在三个广泛应用的多模态动作识别数据集上,PAN-Ensemble与PAN-Unified分别在独立建模与统一建模设置下取得当前最优(SOTA)表现。

原文摘要 · Abstract (English)

While human action recognition has witnessed notable achievements, multimodal methods fusing RGB and skeleton modalities still suffer from their inherent heterogeneity and fail to fully exploit the complementary potential between them. In this paper, we propose PAN, the first human-centric graph representation learning framework for multimodal action recognition, in which token embeddings of RGB patches containing human joints are represented as spatiotemporal graphs. The human-centric graph modeling paradigm suppresses the redundancy in RGB frames and aligns well with skeleton-based methods, thus enabling a more effective and semantically coherent fusion of multimodal features. Since the sampling of token embeddings heavily relies on 2D skeletal data, we further propose attention-based post calibration to reduce the dependency on high-quality skeletal data at a minimal cost interms of model performance. To explore the potential of PAN in integrating with skeleton-based methods, we present two variants: PAN-Ensemble, which employs dual-path graph convolution networks followed by late fusion, and PAN-Unified, which performs unified graph representation learning within a single network. On three widely used multimodal action recognition datasets, both PAN-Ensemble and PAN-Unified achieve state-of-the-art (SOTA) performance in their respective settings of multimodal fusion: separate and unified modeling, respectively.

动作识别图神经网络多模态融合人体姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。