针对行人轨迹预测的多模态难题,提出分模式建模新方法
Where Will They Go? Modelling Multimodal Pedestrian Manoeuvres from Ego-centric Videos

- 将行人过马路与不过马路行为分开建模,提升多模态预测能力
- 在PIE和JAAD数据集上,预测误差比现有方法降低12.3%和9.7%
- 适用于多种框架,可作为通用模块提升轨迹预测性能
基于车载摄像头的行人轨迹预测因涉及行人与车辆、环境的复杂交互及意图不确定性,具有高度多模态特性。现有基于条件变分自编码器(CVAE)的方法通过历史与未来轨迹对齐隐空间,但常抑制多模态性,导致预测结果混合不同行为模式。本文提出MMPM框架,通过行为感知的行人交互模块(PIM)融合视线与手势信息,捕捉行人-车辆与行人-环境交互;并设计基于CVAE的模式感知轨迹预测模块(MTP),分别建模过马路与不过马路两种语义模式。查询式解码器进一步确保解码过程中模式一致性。在PIE和JAAD数据集上的实验表明,该方法优于当前最优基准。所提MTP模块具备模型无关性,可嵌入BiTrap、SGNet等框架中进一步提升性能。此外,我们引入数据驱动的验证协议,从同一场景中检索时空一致的真实轨迹进行对比评估,结果显示帧级位移误差显著优于先前工作。
原文摘要 · Abstract (English)
Pedestrian trajectory prediction from an on-board ego-centric camera is challenging since it depends on complex interactions with vehicles and scene context, as well as the intention of the pedestrian. The task becomes even more challenging since pedestrian intention is often ambiguous from historical observations alone, leading to an inherently multimodal distribution over future trajectories. Existing CVAE-based predictors learn a latent representation by aligning the prior conditioned on trajectory history with the posterior conditioned on both history and future trajectories. This alignment may suppress multimodality in the latent space, leading to mixed-mode predictions that interpolate between distinct future behaviours. In this paper, we propose MMPM, a mode-aware framework that separately models future trajectory distributions into semantically meaningful modes based on the pedestrian's crossing behavior. MMPM consists of two modules: behavior-aware Pedestrian Interaction Module (PIM) that jointly captures pedestrian-vehicle and pedestrian-environment interactions by introducing gaze and hand gestures, and a CVAE-based Mode-aware Trajectory Predictor (MTP) module to model the future trajectory distributions in two modes, crossing and non-crossing the road, separately. A query-based decoder further enforces mode consistency during decoding. Experiments on PIE and JAAD datasets show that our method surpasses state-of-the-art baselines. Our proposed MTP is model-agnostic, which can be integrated into existing frameworks such as BiTrap and SGNet to further improve future trajectory prediction performance. We additionally introduce a data-driven validation protocol that retrieves spatio-temporally consistent real-world trajectories from the same scene and evaluates predictions against these retrieved examples, demonstrating improved frame-wise displacement errors over previous work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。