arXiv:2410.03860cs.CV2024-10CVPR被引 1

用多模态扩散模型预测动作并量化不确定性,提升长期运动精度与人机交互安全。

MDMP: Multi-modal Diffusion for supervised Motion Predictions with uncertainty

  • 融合骨骼数据与文本描述,通过图注意力网络捕捉时空动态。
  • 在长期动作预测上超越现有生成方法,支持多模式输出。
  • 可量化每个关节的置信区域,适合人机协作场景应用。

本文提出一种用于运动预测的多模态扩散模型(MDMP),整合骨骼数据与动作文本描述,生成具有可量化不确定性的长期运动预测结果。现有运动预测或生成方法仅依赖历史运动或文本提示,难以兼顾精度与控制能力,尤其在长时序下表现受限。本方法通过多模态输入增强对人类动作的上下文理解,并采用基于图的变压器框架有效捕捉运动的空间与时间动态。实验表明,该模型在长期运动预测任务中持续优于现有生成技术。此外,利用扩散模型对多模式预测的建模能力,可估计不确定性,通过为各身体关节提供不同置信度的活动区域,显著提升人机交互中的空间感知能力。

原文摘要 · Abstract (English)

This paper introduces a Multi-modal Diffusion model for Motion Prediction (MDMP) that integrates and synchronizes skeletal data and textual descriptions of actions to generate refined long-term motion predictions with quantifiable uncertainty. Existing methods for motion forecasting or motion generation rely solely on either prior motions or text prompts, facing limitations with precision or control, particularly over extended durations. The multi-modal nature of our approach enhances the contextual understanding of human motion, while our graph-based transformer framework effectively capture both spatial and temporal motion dynamics. As a result, our model consistently outperforms existing generative techniques in accurately predicting long-term motions. Additionally, by leveraging diffusion models' ability to capture different modes of prediction, we estimate uncertainty, significantly improving spatial awareness in human-robot interactions by incorporating zones of presence with varying confidence levels for each body joint.

动作预测扩散模型不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。