首次实现人体动作属性的精准操控,可独立调整技术特征而不影响其他动作细节。
Motion Diffusion Autoencoders: Enabling Attribute Manipulation in Human Motion Demonstrated on Karate Techniques
- 提出新型连续旋转姿态表示,解耦骨骼结构与运动轨迹。
- 利用Transformer编码器提取语义嵌入,扩散模型建模随机变化。
- 语义空间线性可操作,适合研究者和动画师精细控制动作特征。
属性操控旨在改变数据点或时间序列的单一属性,同时保持其余部分不变。本文聚焦于人体动作领域,特别是跆拳道动作模式。据我们所知,这是首个成功实现人体动作属性操控的工作。实现该目标的关键在于合适的姿态表示。为此,我们设计了一种新的连续、基于旋转的姿态表示,能够解耦人体骨骼与运动轨迹,同时仍可精确重建原始形态。核心思想是使用Transformer编码器发现高层语义,用扩散概率模型建模剩余随机变化。实验表明,通过Transformer编码器获得的嵌入空间具有语义意义且呈线性,可通过在语义嵌入空间中沿特定方向移动来操控高层属性。所有代码与数据均已公开。
原文摘要 · Abstract (English)
Attribute manipulation deals with the problem of changing individual attributes of a data point or a time series, while leaving all other aspects unaffected. This work focuses on the domain of human motion, more precisely karate movement patterns. To the best of our knowledge, it presents the first success at manipulating attributes of human motion data. One of the key requirements for achieving attribute manipulation on human motion is a suitable pose representation. Therefore, we design a novel continuous, rotation-based pose representation that enables the disentanglement of the human skeleton and the motion trajectory, while still allowing an accurate reconstruction of the original anatomy. The core idea of the manipulation approach is to use a transformer encoder for discovering high-level semantics, and a diffusion probabilistic model for modeling the remaining stochastic variations. We show that the embedding space obtained from the transformer encoder is semantically meaningful and linear. This enables the manipulation of high-level attributes, by discovering their linear direction of change in the semantic embedding space and moving the embedding along said direction. All code and data is made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。