用图嵌入与位置解耦提升行人过街意图预测准确率
GTransPDM: A Graph-embedded Transformer with Positional Decoupling for Pedestrian Crossing Intention Prediction
- 通过位置解耦模块分离横向运动与深度信息
- 在PIE数据集达92%准确率,处理速度仅0.05ms
- 适合自动驾驶中行人行为预测场景
理解并预测行人过街行为意图对自动驾驶车辆安全至关重要。然而,利用图像或环境上下文掩码提取多因素进行时序建模时,常因车载摄像头捕获的行人位置失真,导致预处理误差或效率下降。为此,本文提出GTransPDM——一种带位置解耦模块的图嵌入Transformer,用于行人过街意图预测。首先设计位置解耦模块,将行人横向运动与图像视图中的深度线索分离编码;其次构建图嵌入Transformer,捕捉人体姿态骨架的时空动态,融合位置、骨架及自车运动等关键因素。实验表明,该方法在PIE数据集上达到92%准确率,在JAAD数据集上达87%,处理速度仅为0.05ms,优于现有最先进方法。
原文摘要 · Abstract (English)
Understanding and predicting pedestrian crossing behavioral intention is crucial for the driving safety of autonomous vehicles. Nonetheless, challenges emerge when using promising images or environmental context masks to extract various factors for time-series network modeling, causing pre-processing errors or a loss of efficiency. Typically, pedestrian positions captured by onboard cameras are often distorted and do not accurately reflect their actual movements. To address these issues, GTransPDM -- a Graph-embedded Transformer with a Position Decoupling Module -- was developed for pedestrian crossing intention prediction by leveraging multi-modal features. First, a positional decoupling module was proposed to decompose pedestrian lateral motion and encode depth cues in the image view. Then, a graph-embedded Transformer was designed to capture the spatio-temporal dynamics of human pose skeletons, integrating essential factors such as position, skeleton, and ego-vehicle motion. Experimental results indicate that the proposed method achieves 92% accuracy on the PIE dataset and 87% accuracy on the JAAD dataset, with a processing speed of 0.05ms. It outperforms the state-of-the-art in comparison.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。