用图像关键帧控制动作生成,提升精度与用户意图契合度
IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model
- 分离轨迹与姿态输入,分阶段优化并行编码
- 在HumanML3D和KIT-ML数据集上全指标超越现有方法
- 结合多模态大模型预处理,更贴合用户文本与图像输入
现有基于轨迹和姿态输入的人体动作生成方法对两种模态进行全局处理,导致输出效果不佳。本文提出IKMo,一种基于扩散模型的图像关键帧动作生成方法,实现轨迹与姿态的解耦。输入通过两阶段条件框架处理:第一阶段使用专用优化模块精炼输入;第二阶段由轨迹编码器与姿态编码器并行编码。随后,融合后的轨迹与姿态数据由运动ControlNet引导,生成高空间与语义保真度的动作。在HumanML3D和KIT-ML数据集上的实验表明,该方法在轨迹-关键帧约束下所有指标均优于当前最先进方法。此外,引入基于MLLM的智能代理预处理输入:根据用户提供的文本和关键帧图像,自动提取动作描述、关键帧姿态与轨迹,并作为优化输入送入生成模型。10名参与者的用户研究证实,该预处理显著提升生成动作与用户期望的一致性。我们认为,该方法通过扩散模型提升了动作生成的保真度与可控性。
原文摘要 · Abstract (English)
Existing human motion generation methods with trajectory and pose inputs operate global processing on both modalities, leading to suboptimal outputs. In this paper, we propose IKMo, an image-keyframed motion generation method based on the diffusion model with trajectory and pose being decoupled. The trajectory and pose inputs go through a two-stage conditioning framework. In the first stage, the dedicated optimization module is applied to refine inputs. In the second stage, trajectory and pose are encoded via a Trajectory Encoder and a Pose Encoder in parallel. Then, motion with high spatial and semantic fidelity is guided by a motion ControlNet, which processes the fused trajectory and pose data. Experiment results based on HumanML3D and KIT-ML datasets demonstrate that the proposed method outperforms state-of-the-art on all metrics under trajectory-keyframe constraints. In addition, MLLM-based agents are implemented to pre-process model inputs. Given texts and keyframe images from users, the agents extract motion descriptions, keyframe poses, and trajectories as the optimized inputs into the motion generation model. We conducts a user study with 10 participants. The experiment results prove that the MLLM-based agents pre-processing makes generated motion more in line with users' expectation. We believe that the proposed method improves both the fidelity and controllability of motion generation by the diffusion model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。