用人体姿态提升轨迹预测,让自动驾驶更懂人的动作意图。
Social-Pose: Enhancing Trajectory Prediction with Human Body Pose
- 基于注意力机制的姿势编码器,捕捉人群姿态与社交关系。
- 在多个数据集上优于传统方法,最高提升12.3%的预测精度。
- 适用于机器人导航,对噪声姿态有良好鲁棒性。
准确的人体轨迹预测是自动驾驶安全的关键任务,但现有模型未能充分利用人类在空间移动时无意识传递的视觉线索。本文提出「Social-Pose」,通过人体姿态而非仅坐标位置来预测轨迹。设计了一种基于注意力的姿势编码器,有效捕获场景中所有行人的姿态及其社会关系,并可嵌入多种预测架构。在联合跟踪自动(Joint Track Auto)等合成数据集,以及真实数据集(Human3.6M、Pedestrians and Cyclists in Road Traffic、JRDB)上进行了大量实验,结果表明在基于LSTM、GAN、MLP和Transformer的主流模型上均取得显著提升。还对比了2D与3D姿态的效果,分析了噪声姿态的影响,并验证了其在机器人导航中的适用性。
原文摘要 · Abstract (English)
Accurate human trajectory prediction is one of the most crucial tasks for autonomous driving, ensuring its safety. Yet, existing models often fail to fully leverage the visual cues that humans subconsciously communicate when navigating the space. In this work, we study the benefits of predicting human trajectories using human body poses instead of solely their Cartesian space locations in time. We propose `Social-pose', an attention-based pose encoder that effectively captures the poses of all humans in a scene and their social relations. Our method can be integrated into various trajectory prediction architectures. We have conducted extensive experiments on state-of-the-art models (based on LSTM, GAN, MLP, and Transformer), and showed improvements over all of them on synthetic (Joint Track Auto) and real (Human3.6M, Pedestrians and Cyclists in Road Traffic, and JRDB) datasets. We also explored the advantages of using 2D versus 3D poses, as well as the effect of noisy poses and the application of our pose-based predictor in robot navigation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。