建模行人与车辆3D互动,提升自动驾驶预测精度
Modeling 3D Pedestrian-Vehicle Interactions for Vehicle-Conditioned Pose Forecasting
- 用车辆信息引导行人姿态预测,引入交叉注意力融合特征
- 在增强版Waymo-3DSkelMo数据集上,预测误差显著降低
- 适合研究自动驾驶交互建模的工程师和研究人员
准确预测行人运动对复杂城市环境中的自动驾驶安全至关重要。本文提出一种3D车辆条件化行人姿态预测框架,显式引入周围车辆信息。为此,我们对Waymo-3DSkelMo数据集进行了增强,添加了对齐的3D车辆边界框,以实现多智能体行人-车辆交互的真实建模。设计了一种基于行人与车辆数量的采样方案,支持不同交互复杂度下的训练。所提网络在TBIFormer基础上加入专用车辆编码器和行人-车辆交互交叉注意力模块,融合行人与车辆特征,使预测同时依赖历史行人运动和周边车辆状态。大量实验表明,该方法显著提升了预测准确性,验证了多种行人-车辆交互建模方式的有效性,凸显了车辆感知的3D姿态预测对自动驾驶的重要性。代码已开源。
原文摘要 · Abstract (English)
Accurately predicting pedestrian motion is crucial for safe and reliable autonomous driving in complex urban environments. In this work, we present a 3D vehicle-conditioned pedestrian pose forecasting framework that explicitly incorporates surrounding vehicle information. To support this, we enhance the Waymo-3DSkelMo dataset with aligned 3D vehicle bounding boxes, enabling realistic modeling of multi-agent pedestrian-vehicle interactions. We introduce a sampling scheme to categorize scenes by pedestrian and vehicle count, facilitating training across varying interaction complexities. Our proposed network adapts the TBIFormer architecture with a dedicated vehicle encoder and pedestrian-vehicle interaction cross-attention module to fuse pedestrian and vehicle features, allowing predictions to be conditioned on both historical pedestrian motion and surrounding vehicles. Extensive experiments demonstrate substantial improvements in forecasting accuracy and validate different approaches for modeling pedestrian-vehicle interactions, highlighting the importance of vehicle-aware 3D pose prediction for autonomous driving. Code is available at: https://github.com/GuangxunZhu/VehCondPose3D
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。