通过模态耦合提升扩散模型运动生成的可控性与真实性
Controllable Motion Generation via Diffusion Modal Coupling
- 利用多模态先验和强模态耦合,从不同行为模式启动去噪过程
- 在Waymo数据集和Maze2D环境中,生成结果更真实、多样且可控
- 无需显式条件输入,仍能精准控制生成轨迹,适合机器人任务
扩散模型因其生成系统状态与行为多模态分布的能力,在机器人领域受到广泛关注。然而,如何在不牺牲真实性的前提下实现精确控制仍是关键挑战,尤其在运动规划或轨迹预测等需满足物理约束与任务目标的应用中。本文提出一种新框架,通过引入多模态先验并强化模态耦合,使去噪过程可直接从对应不同系统行为的先验模式启动,确保采样结果与训练分布一致。我们在Waymo数据集上评估运动预测性能,并在Maze2D环境中测试多任务控制。实验表明,该方法优于基于引导的技术和单模态先验的条件模型,在无显式条件输入的情况下仍能实现更高的保真度、多样性与可控性。整体上,为机器人中的可控运动生成提供了更可靠、可扩展的解决方案。
原文摘要 · Abstract (English)
Diffusion models have recently gained significant attention in robotics due to their ability to generate multi-modal distributions of system states and behaviors. However, a key challenge remains: ensuring precise control over the generated outcomes without compromising realism. This is crucial for applications such as motion planning or trajectory forecasting, where adherence to physical constraints and task-specific objectives is essential. We propose a novel framework that enhances controllability in diffusion models by leveraging multi-modal prior distributions and enforcing strong modal coupling. This allows us to initiate the denoising process directly from distinct prior modes that correspond to different possible system behaviors, ensuring sampling to align with the training distribution. We evaluate our approach on motion prediction using the Waymo dataset and multi-task control in Maze2D environments. Experimental results show that our framework outperforms both guidance-based techniques and conditioned models with unimodal priors, achieving superior fidelity, diversity, and controllability, even in the absence of explicit conditioning. Overall, our approach provides a more reliable and scalable solution for controllable motion generation in robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。