通过骨架运动建模提升动作识别性能,支持少样本与零样本迁移。
T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

- 基于骨架序列与多模态对比学习,显式建模人体运动特征。
- 在多个数据集上表现优于现有方法,零样本设置下仍具强泛化能力。
- 适用于需要理解人体动态行为的场景,如智能监控与人机交互。
视觉语言模型(如CLIP)在多种视觉理解任务中表现出色,但多数模型依赖图像或视频的外观监督,未显式建模人体运动,而运动是细粒度、以人为中心的动作识别的关键。为此,我们提出可迁移的骨架运动表征(T-MOR),利用视频和语言监督从骨架序列中学习可迁移的动作表示。T-MOR采用多模态对比学习策略,对齐骨架运动与视觉、文本表征,推理时仅需轻量级骨架输入。为支持大规模预训练,我们构建了PoseCap-1M数据集,包含超过一百万组同步的视频、骨架和文本三元组,涵盖多样人体活动。我们在多个以人为核心的动作识别基准上评估T-MOR,包括动作分类与帧级时间检测。实验表明,T-MOR在Toyota Smarthome、Penn Action、UAV-Human、TSU和Charades等多个数据集上持续提升性能;同时在少样本与零样本设置下展现强大泛化能力,验证了以运动为中心、具身化表征在可迁移动作理解中的有效性。
原文摘要 · Abstract (English)
Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing models rely primarily on appearance-level supervision from images or videos, and do not explicitly model human motion, which is essential for fine-grained and human-centric action recognition task as actions are defined by temporally structured and physically grounded body movements. To address this problem, we propose Transferable skeleton MOtion Representation (T-MOR), a motion-aware framework that learns transferable action representations from skeleton sequences with the aid of video and language supervision during training. T-MOR adopts a multi-modal contrastive learning scheme that aligns skeleton motion with visual and textual representations, while performing inference using only lightweight skeleton inputs. To support large-scale pre-training, we construct PoseCap-1M, a new dataset that contains over one million synchronized video, skeleton, and text triplets covering diverse human activities. We evaluate T-MOR on a range of human-centric action recognition benchmarks, including action classification and frame-wise temporal detection. Experimental results show that T-MOR consistently improves performance across multiple datasets, such as Toyota Smarthome, Penn Action, UAV-Human, TSU, and Charades. In addition, T-MOR demonstrates strong generalization ability in few-shot and zero-shot settings, highlighting the effectiveness of motion-centric and embodied representations for transferable action understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。