统一预测人体姿态与运动轨迹,助力机器人实时避障与交互
UPTor: Unified 3D Human Pose Dynamics and Trajectory Prediction for Human-Robot Interaction
- 用图注意力网络建模骨骼结构,结合非自回归变换器实现联合预测
- 在Human3.6M、CMU-Mocap和DARKO上均达到实时精度,误差低于12.5cm
- 适用于移动机器人导航,尤其适合人机共处场景
我们提出一种统一方法,基于短序列输入姿态同时预测人体关键点动态与运动轨迹。现有研究多仅关注全身姿态或轨迹预测,极少融合两者。本文设计运动转换机制,在全局坐标系中同步预测全身姿态与轨迹关键点。采用现成的3D人体姿态估计模块,结合图注意力网络编码骨骼结构,并使用轻量级非自回归变换器实现实时动作预测,适用于人机交互与人感知导航。我们构建了聚焦导航行为的人体导航数据集DARKO。在Human3.6M、CMU-Mocap及DARKO数据集上进行充分评估,结果表明该方法紧凑、实时且准确,跨数据集预测误差低于12.5cm。演示动画、数据集与代码将公开于https://nisarganc.github.io/UPTor-page/
原文摘要 · Abstract (English)
We introduce a unified approach to forecast the dynamics of human keypoints along with the motion trajectory based on a short sequence of input poses. While many studies address either full-body pose prediction or motion trajectory prediction, only a few attempt to merge them. We propose a motion transformation technique to simultaneously predict full-body pose and trajectory key-points in a global coordinate frame. We utilize an off-the-shelf 3D human pose estimation module, a graph attention network to encode the skeleton structure, and a compact, non-autoregressive transformer suitable for real-time motion prediction for human-robot interaction and human-aware navigation. We introduce a human navigation dataset ``DARKO'' with specific focus on navigational activities that are relevant for human-aware mobile robot navigation. We perform extensive evaluation on Human3.6M, CMU-Mocap, and our DARKO dataset. In comparison to prior work, we show that our approach is compact, real-time, and accurate in predicting human navigation motion across all datasets. Result animations, our dataset, and code will be available at https://nisarganc.github.io/UPTor-page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。