H-MoRe通过自监督学习提取人体运动细节,提升动作分析性能。
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
- 基于人体运动学设计世界-局部流矩阵,动态保留有效动作。
- 在步态识别、动作识别等任务上显著提升,最高增益达16.01%。
- 适合实时应用,推理速度达34帧/秒,开源可复现。
本文提出H-MoRe,一种新颖的自监督人体中心运动表征学习框架。该方法不依赖合成数据,直接从真实场景中学习,融合人体姿态与体形信息,动态保留相关人体运动并过滤背景干扰。受运动学启发,H-MoRe以矩阵形式表示每个身体点的绝对与相对运动,称为世界-局部流(world-local flows),精准捕捉细微运动特征。实验表明,该方法在多个下游任务中表现优异:步态识别(CL@R1: +16.01%)、动作识别(Acc@1: +8.92%)、视频生成(FVD: -67.07%)。同时具备高推理效率(34 fps),适用于大多数实时场景。模型与代码将在论文发表后公开。
原文摘要 · Abstract (English)
In this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filtering out background movement. Notably, unlike previous methods relying on fully supervised learning from synthetic data, H-MoRe learns directly from real-world scenarios in a self-supervised manner, incorporating both human pose and body shape information. Inspired by kinematics, H-MoRe represents absolute and relative movements of each body point in a matrix format that captures nuanced motion details, termed world-local flows. H-MoRe offers refined insights into human motion, which can be integrated seamlessly into various action-related applications. Experimental results demonstrate that H-MoRe brings substantial improvements across various downstream tasks, including gait recognition(CL@R1: +16.01%), action recognition(Acc@1: +8.92%), and video generation(FVD: -67.07%). Additionally, H-MoRe exhibits high inference efficiency (34 fps), making it suitable for most real-time scenarios. Models and code will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。