提出新模型,让骨骼动作生成更聚焦真实运动形态。
An Elastic Shape Variational Autoencoder for Skeleton Pose Trajectories

- 用几何流形方法剥离视角、尺度等干扰因素
- 在步态分析和动作识别任务中均优于主流模型
- 适合研究人体运动建模与临床评估的学者
深度生成模型虽能处理复杂数据,但标准变分自编码器(VAE)在处理人体骨骼序列时,常将大量容量用于相机角度、缩放、视角和执行速度等无关因素。本文提出弹性形状变分自编码器(ES-VAE),基于肯德尔形状流形上的传输平方根速度场(TSRVF)表示,自动消除刚性平移、旋转和全局缩放,以及时间速率变化,从而分离出底层形状动态。编码器利用黎曼对数映射将骨骼序列映射至低维潜在空间,解码器通过指数映射重建序列。在两个数据集上验证:一是分析步态周期以预测临床移动评分并分类健康与中风人群;二是在NTU RGB+D数据集上进行动作识别。结果表明,ES-VAE持续优于标准VAE及多种序列建模范式,包括时序卷积网络、Transformer和图卷积网络。该方法为姿态形状流形上的纵向数据生成建模提供了一种严谨框架,相比现有深度学习方法,显著提升了潜在表示质量和下游性能。
原文摘要 · Abstract (English)
Deep generative models provide flexible frameworks for modeling complex, structured data such as images, videos, 3D objects, and texts. However, when applied to sequences of human skeletons, standard variational autoencoders (VAEs) often allocate substantial capacity to nuisance factors-such as camera orientation, subject scale, viewpoint, and execution speed-rather than the intrinsic geometry of shapes and their motion. We propose the Elastic Shape - Variational Autoencoder (ES-VAE), a geometry-aware generative model for skeletal trajectories that leverages the transported square-root velocity field (TSRVF) representation on Kendall's shape manifold. This representation inherently removes rigid translations, rotations, and global scaling of shapes, and temporal rate variability of sequences, isolating the underlying shape dynamics. The ES-VAE encoder maps skeletal sequences to a low-dimensional latent space incorporating the Riemannian logarithm map, while the decoder reconstructs sequences using the corresponding exponential map. We demonstrate the effectiveness of ES-VAE on two datasets. First, we analyze skeletal gait cycles to predict clinical mobility scores and classify subjects into healthy and post-stroke groups. Second, we evaluate action recognition on the NTU RGB+D dataset. Across both settings, ES-VAE consistently outperforms standard VAEs and a range of sequence modeling baselines, including temporal convolutional networks, transformers, and graph convolutional networks. More broadly, ES-VAE provides a principled framework for learning generative models of longitudinal data on pose shape manifolds, offering improved latent representation and downstream performance compared to existing deep learning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。