用全局绝对坐标生成动作,效果更好且更易扩展。
Absolute Coordinates Make Motion Generation Easy
- 采用全局空间中的关节绝对坐标,取代传统相对编码。
- 在简单Transformer上实现更高动作保真度和文本对齐度。
- 天然支持动作编辑与控制,无需额外设计或训练。
当前最先进的文本到动作生成模型依赖于由HumanML3D推广的、基于骨盆和前一帧的局部相对运动表示,该表示虽简化了早期模型的训练,但对扩散模型引入关键限制,并阻碍下游任务应用。本文重新审视运动表示,提出一种被长期弃用却极为简化的替代方案:全局空间中的关节绝对坐标。通过系统性分析,我们证明该方法在仅使用简单Transformer主干且无辅助运动感知损失的情况下,显著提升动作保真度、改善文本对齐,并具备强可扩展性。此外,该形式自然支持文本驱动的动作控制及时间/空间编辑,无需额外的任务特定重构或昂贵的控制信号分类器引导生成。最后,我们展示了直接从文本生成SMPL-H网格顶点序列的潜力,为未来运动相关研究与应用奠定坚实基础。
原文摘要 · Abstract (English)
State-of-the-art text-to-motion generation models rely on the kinematic-aware, local-relative motion representation popularized by HumanML3D, which encodes motion relative to the pelvis and to the previous frame with built-in redundancy. While this design simplifies training for earlier generation models, it introduces critical limitations for diffusion models and hinders applicability to downstream tasks. In this work, we revisit the motion representation and propose a radically simplified and long-abandoned alternative for text-to-motion generation: absolute joint coordinates in global space. Through systematic analysis of design choices, we show that this formulation achieves significantly higher motion fidelity, improved text alignment, and strong scalability, even with a simple Transformer backbone and no auxiliary kinematic-aware losses. Moreover, our formulation naturally supports downstream tasks such as text-driven motion control and temporal/spatial editing without additional task-specific reengineering and costly classifier guidance generation from control signals. Finally, we demonstrate promising generalization to directly generate SMPL-H mesh vertices in motion from text, laying a strong foundation for future research and motion-related applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。