arXiv:2605.13729cs.CVcs.AI2026-05被引 1

解决文本与轨迹冲突,实现精准动作生成

Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation

论文配图:Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
图 1 · 摘自论文原文
  • 分阶段生成:先控轨迹再补全动作,避免条件冲突
  • 在HumanML3D和KIT数据集上达到最优控制精度与动作质量
  • 适合需要精确轨迹控制的动作生成任务

轨迹控制的人体动作生成旨在根据文本描述和空间轨迹生成真实动作。现有方法存在两大问题:一是文本与轨迹条件冲突干扰去噪过程,导致动作质量下降或轨迹偏离;二是冗余运动表示引发动作组件不一致,造成轨迹控制不稳定。为此,我们提出CMC框架,采用解耦的分治策略,包含轨迹控制与动作补全两个级联阶段。第一阶段使用扩散模型在轨迹引导下生成受控关节的简化表示,确保轨迹跟踪准确稳定;第二阶段通过文本条件扩散修复模型,以第一阶段输出为部分观测,生成完整人体动作。为缓解有限修复训练数据带来的过拟合,引入选择性修复机制(SIM),在训练中交替进行文本到动作生成与动作修复任务。在HumanML3D和KIT数据集上的实验表明,CMC在控制精度与动作质量方面均达到当前最优水平,验证了其在多模态条件与表示协调上的有效性。

原文摘要 · Abstract (English)

Trajectory-controlled human motion generation aims to synthesize realistic human motions conditioned on both textual descriptions and spatial trajectories. However, existing methods suffer from two critical limitations: first, the conflict between text and trajectory conditions disrupts the denoising process, resulting in compromised motion quality or inaccurate trajectory following; second, the use of redundant motion representations introduces inconsistencies between motion components, leading to instability during trajectory control. To address these challenges, we propose CMC, a decoupled framework that effectively coordinates text and trajectory conditions through a divide-and-conquer strategy. CMC follows a divide-and-conquer paradigm, comprising two cascaded stages: Trajectory Control and Motion Completion. In the first stage, a diffusion model generates a simplified representation of the controlled joints under trajectory guidance, based on the given trajectories, ensuring accurate and stable trajectory following. In the second stage, a text-conditioned diffusion inpainting model generates full-body motions using the simplified representation from the first stage as partial observations. To mitigate overfitting caused by limited inpainting training data, we further introduce the Selective Inpainting Mechanism (SIM), which alternates between text-to-motion generation and motion inpainting tasks during training. Experiments on HumanML3D and KIT datasets demonstrate that CMC achieves state-of-the-art performance in control accuracy and motion quality, demonstrating its effectiveness in coordinating multimodal conditions and representations.

动作生成轨迹控制扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。