arXiv:2603.27103cs.CV2026-03被引 1

用骨骼+语言联合建模,提升动作识别的语义理解能力

LLM Enhanced Action Recognition via Hierarchical Global-Local Skeleton-Language Model

  • 分层全局-局部网络捕捉骨骼的长程依赖与时空关系
  • 结合视觉语言模型生成动作描述,增强语义表征
  • 适合需要细粒度动作理解的场景,如智能监控、医疗分析

基于骨骼的人体动作识别近年取得显著进展,但现有GCN方法多依赖短程运动拓扑,难以捕捉长程关节依赖与复杂时序动态,且因动作语义建模不足,限制了跨模态对齐与理解。为此,本文提出分层全局-局部骨骼-语言模型(HocSLM),首先设计分层全局-局部网络(HGLNet),包含复合拓扑空间模块与双路径分层时序模块,通过多层级全局与局部模块协同建模,动态融合不同尺度信息,同时保留人体物理结构先验知识,显著增强复杂时空关系表征能力;其次,利用大视觉语言模型(VLM)对原始RGB视频生成文本描述,为骨架-语言模型提供丰富动作语义;最后,引入骨架-语言序列融合模块,结合HGLNet特征与生成描述,通过骨架-语言模型(SLM)在统一语义空间中精确对齐骨骼时空特征与文本动作描述,显著提升模型语义区分能力与跨模态理解能力。大量实验表明,所提HocSLM在三个主流基准数据集NTU RGB+D 60、NTU RGB+D 120和Northwestern-UCLA上均达到当前最优性能。

原文摘要 · Abstract (English)

Skeleton-based human action recognition has achieved remarkable progress in recent years. However, most existing GCN-based methods rely on short-range motion topologies, which not only struggle to capture long-range joint dependencies and complex temporal dynamics but also limit cross-modal semantic alignment and understanding due to insufficient modeling of action semantics. To address these challenges, we propose a hierarchical global-local skeleton-language model (HocSLM), enabling the large action model be more representative of action semantics. First, we design a hierarchical global-local network (HGLNet) that consists of a composite-topology spatial module and a dual-path hierarchical temporal module. By synergistically integrating multi-level global and local modules, HGLNet achieves dynamically collaborative modeling at both global and local scales while preserving prior knowledge of human physical structure, significantly enhancing the model's representation of complex spatio-temporal relationships. Then, a large vision-language model (VLM) is employed to generate textual descriptions by passing the original RGB video sequences to this model, providing the rich action semantics for further training the skeleton-language model. Furthermore, we introduce a skeleton-language sequential fusion module by combining the features from HGLNet and the generated descriptions, which utilizes a skeleton-language model (SLM) for aligning skeletal spatio-temporal features and textual action descriptions precisely within a unified semantic space. The SLM model could significantly enhance the HGLNet's semantic discrimination capabilities and cross-modal understanding abilities. Extensive experiments demonstrate that the proposed HocSLM achieves the state-of-the-art performance on three mainstream benchmark datasets: NTU RGB+D 60, NTU RGB+D 120, and Northwestern-UCLA.

动作识别骨骼建模多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。