用分层运动描述提升动作理解与生成精度
KinMo: Kinematic-aware Human Motion Understanding and Generation
- 构建分层运动表示,融合肢体动态与交互关系
- 文本-动作检索准确率显著提升,支持精细动作编辑
- 适合需要高精度动作生成的研究者与开发者
当前人体动作合成框架依赖全局动作描述,导致语义与动作模态间存在鸿沟。单一粗粒度描述(如“跑”)无法捕捉速度、肢体姿态和运动动力学等细节,造成歧义。为此,我们提出KinMo,一种基于分层可描述运动表示的统一框架,扩展了全局动作,引入肢体组运动及其相互作用。设计自动化标注流程,生成高质量细粒度描述,形成KinMo数据集,提供可扩展且低成本的数据增强方案。为利用结构化描述,提出分层文本-动作对齐机制,逐步整合运动细节,提升语义理解能力。进一步提出粗到细的动作生成流程,借助增强的空间理解改善动作合成效果。实验表明,KinMo显著提升动作理解能力,体现在文本-动作检索性能提升,并实现更精细的动作生成与编辑。项目页面:https://andypinxinliu.github.io/KinMo
原文摘要 · Abstract (English)
Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as run, fails to capture details such as variations in speed, limb positioning, and kinematic dynamics, leading to ambiguities between text and motion modalities. To address this challenge, we introduce KinMo, a unified framework built on a hierarchical describable motion representation that extends beyond global actions by incorporating kinematic group movements and their interactions. We design an automated annotation pipeline to generate high-quality, fine-grained descriptions for this decomposition, resulting in the KinMo dataset and offering a scalable and cost-efficient solution for dataset enrichment. To leverage these structured descriptions, we propose Hierarchical Text-Motion Alignment that progressively integrates additional motion details, thereby improving semantic motion understanding. Furthermore, we introduce a coarse-to-fine motion generation procedure to leverage enhanced spatial understanding to improve motion synthesis. Experimental results show that KinMo significantly improves motion understanding, demonstrated by enhanced text-motion retrieval performance and enabling more fine-grained motion generation and editing capabilities. Project Page: https://andypinxinliu.github.io/KinMo
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。