arXiv:2603.15546cs.CVcs.GR2026-03被引 27

用700小时动捕数据训练,实现高保真可控人体运动生成

Kimodo: Scaling Controllable Human Motion Generation

  • 基于分解根部与躯干预测的两阶段去噪架构
  • 支持文本和多种骨骼约束条件,生成运动质量高
  • 适合需要精细控制的人体动作生成场景

高质量人体动作数据在机器人、仿真和娱乐领域日益重要。现有生成模型虽可通过文本提示或姿态约束生成动作,但受限于公开动捕数据集规模小,导致动作质量、控制精度与泛化能力不足。本文提出Kimodo,一个在700小时光学动捕数据上训练的可表达、可控制的运动扩散模型。该模型通过精心设计的动作表示与两阶段去噪器架构,支持全身体关键帧、稀疏关节位置/旋转、2D路径点及密集2D轨迹等多种约束条件,同时有效减少运动伪影。实验验证了关键设计选择,并分析了数据量与模型规模对性能的影响。

原文摘要 · Abstract (English)

High-quality human motion data is becoming increasingly important for applications in robotics, simulation, and entertainment. Recent generative models offer a potential data source, enabling human motion synthesis through intuitive inputs like text prompts or kinematic constraints on poses. However, the small scale of public mocap datasets has limited the motion quality, control accuracy, and generalization of these models. In this work, we introduce Kimodo, an expressive and controllable kinematic motion diffusion model trained on 700 hours of optical motion capture data. Our model generates high-quality motions while being easily controlled through text and a comprehensive suite of kinematic constraints including full-body keyframes, sparse joint positions/rotations, 2D waypoints, and dense 2D paths. This is enabled through a carefully designed motion representation and two-stage denoiser architecture that decomposes root and body prediction to minimize motion artifacts while allowing for flexible constraint conditioning. Experiments on the large-scale mocap dataset justify key design decisions and analyze how the scaling of dataset size and model size affect performance.

人体运动生成扩散模型动作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。