用双曲空间建模骨骼动作层次结构,提升少样本动作识别性能
SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

- 采用双曲空间嵌入骨骼序列与语义,天然捕捉人体层级结构
- 在三个数据集上超越当前最佳方法,提升显著
- 无需训练的投票缓存机制,适合资源受限场景
基于骨骼的动作识别旨在从关节序列中理解人类行为,尤其在少样本设置下挑战巨大——每个新动作仅有一个标注样本。核心难点在于学习能捕捉人体运动层次与组合结构、并在极端数据稀缺下有效对齐高层语义的表征。现有方法多依赖欧氏嵌入与低层运动线索,难以建模骨骼数据的树状组织,限制了跨模态对齐与泛化能力。本文提出SkelHCC,一种统一的骨架双曲CLIP驱动缓存适配框架。SkelHCC引入显式分层双曲CLIP(EH-HCLIP)模块,将骨架序列与动作语言嵌入共享双曲空间。利用双曲几何的负曲率与指数体积增长特性,自然编码人体关节-部位-躯干的层级关系,生成结构一致的跨模态表示。为支持高效少样本适应,进一步集成免训练的LLM引导多粒度投票缓存(LMV-Cache),实现上下文感知推理。在NTU RGB+D 60、NTU RGB+D 120和PKU-MMD数据集上的实验表明,SkelHCC持续优于现有最先进方法。
原文摘要 · Abstract (English)
Skeleton-based action recognition aims to understand human behaviors from body joint sequences and is especially challenging in the one-shot setting, where only a single labeled exemplar is available for each novel action. A key challenge is learning representations that capture the hierarchical and compositional structure of human motion while aligning effectively with high-level action semantics under extreme data scarcity. Existing approaches, largely based on Euclidean embeddings and low-level motion cues, struggle to model the tree-like organization of skeleton data, limiting cross-modal alignment and generalization to unseen action categories. We propose SkelHCC, a unified skeleton hyperbolic CLIP-driven cache adaptation framework for one-shot skeleton-based action recognition. SkelHCC introduces an Explicitly Hierarchical Hyperbolic CLIP (EH-HCLIP) module that embeds skeleton sequences and action language into a shared hyperbolic space. By leveraging the negative curvature and exponential volume growth of hyperbolic geometry, EH-HCLIP naturally encodes the joint-part-body hierarchy of human anatomy and yields structurally consistent cross-modal representations. To support efficient one-shot adaptation, SkelHCC further integrates a training-free LLM-guided Multi-granularity Voting Cache (LMV-Cache) for context-aware inference. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD demonstrate that SkelHCC consistently outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。