看人走路的腿比看上半身更能预测其动向,适合做机器人导航。
Legs Over Arms: On the Predictive Value of Lower-Body Pose for Human Trajectory Prediction from Egocentric Robot Perception
- 用人体下肢3D关键点替代传统点质量模型,提升轨迹预测精度。
- 在JRDB和全景视频新数据集上,平均位移误差降低13%。
- 单目全景视觉也能捕捉有效运动线索,利于低成本感知设计。
在拥挤环境中,预测人类轨迹对社交机器人导航至关重要。现有方法多将人视为点质量,本文系统评估了2D与3D骨骼关键点及衍生生物力学特征作为额外输入的预测效用。基于JRDB数据集及一个新的包含360°全景视频的社会导航数据集,研究发现关注下肢3D关键点可使平均位移误差降低13%;进一步融合生物力学特征,误差再降1-4%。值得注意的是,即使使用从等距投影全景图像中提取的2D关键点,性能提升依然显著,表明单目环视视觉即可捕获有效的运动预测线索。结果表明,机器人通过观察人的腿部即可高效预测其移动路径,为社交机器人感知系统设计提供了可操作的指导。
原文摘要 · Abstract (English)
Predicting human trajectory is crucial for social robot navigation in crowded environments. While most existing approaches treat human as point mass, we present a study on multi-agent trajectory prediction that leverages different human skeletal features for improved forecast accuracy. In particular, we systematically evaluate the predictive utility of 2D and 3D skeletal keypoints and derived biomechanical cues as additional inputs. Through a comprehensive study on the JRDB dataset and another new dataset for social navigation with 360-degree panoramic videos, we find that focusing on lower-body 3D keypoints yields a 13% reduction in Average Displacement Error and augmenting 3D keypoint inputs with corresponding biomechanical cues provides a further 1-4% improvement. Notably, the performance gain persists when using 2D keypoint inputs extracted from equirectangular panoramic images, indicating that monocular surround vision can capture informative cues for motion forecasting. Our finding that robots can forecast human movement efficiently by watching their legs provides actionable insights for designing sensing capabilities for social robot navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。