用视觉相机运动信息提升行人轨迹预测准确率
TrajMamba: An Ego-Motion-Guided Mamba Model for Pedestrian Trajectory Prediction from an Egocentric Perspective
- 双Mamba编码器分别提取行人与相机运动特征
- 结合相机运动引导,相对运动建模提升预测精度
- 在PIE和JAAD数据集上达到领先效果
从第一人称视角进行行人未来轨迹预测是自动驾驶与机器人导航中的关键任务。其难点在于摄像头与行人间复杂的动态相对运动。为此,我们提出一种基于Mamba模型的相机运动引导轨迹预测网络。首先,使用两个Mamba模型作为编码器,分别从行人运动和摄像头运动中提取行人运动特征与相机运动特征。然后,设计一个相机运动引导的Mamba解码器,通过将行人运动特征作为历史上下文、相机运动特征作为引导信号,显式建模行人与车辆间的相对运动,以捕获未来时刻的解码特征。最终,从对应未来时间戳的解码特征生成预测轨迹。大量实验表明,所提模型在PIE与JAAD数据集上均取得当前最优性能。
原文摘要 · Abstract (English)
Future trajectory prediction of a tracked pedestrian from an egocentric perspective is a key task in areas such as autonomous driving and robot navigation. The challenge of this task lies in the complex dynamic relative motion between the ego-camera and the tracked pedestrian. To address this challenge, we propose an ego-motion-guided trajectory prediction network based on the Mamba model. Firstly, two Mamba models are used as encoders to extract pedestrian motion and ego-motion features from pedestrian movement and ego-vehicle movement, respectively. Then, an ego-motion guided Mamba decoder that explicitly models the relative motion between the pedestrian and the vehicle by integrating pedestrian motion features as historical context with ego-motion features as guiding cues to capture decoded features. Finally, the future trajectory is generated from the decoded features corresponding to the future timestamps. Extensive experiments demonstrate the effectiveness of the proposed model, which achieves state-of-the-art performance on the PIE and JAAD datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。