分三步建模行人轨迹,提升预测精度与可解释性
Three-Step Hierarchical Transformer for Multi-Pedestrian Trajectory Prediction

- 分阶段处理时序、多模态与社交交互,结构清晰
- 在JRDB和Urban数据集上达到当前最佳性能
- 适合需要精准预测复杂行为的智能交通场景
行人轨迹预测需建模时间动态、多模态线索及拥挤环境中的社交互动。现有方法常分别处理或在昂贵注意力模块中混杂这些因素,限制了可扩展性、灵活性与可解释性。本文提出三步分层Transformer,明确分离时序编码、多模态融合与场景级交互推理。轻量GRU摘要实现高效跨模态注意力,随时间变化的行人令牌社交注意力以可控代价捕捉行人间影响。在JTA、JRDB及行人与骑行者道路交通数据集上的实验表明,该模型在真实世界数据集(JRDB、Urban)上达到领先性能,在JTA上表现具有竞争力。消融与定性分析验证了各阶段贡献,并展示了模型对提前转向等复杂行为的预测能力。
原文摘要 · Abstract (English)
Pedestrian trajectory prediction requires modeling temporal dynamics, multimodal cues, and social interactions in crowded environments. Existing methods often address these factors separately or entangle them in costly attention blocks, limiting scalability, flexibility, and interpretability. We propose a three-step hierarchical Transformer that explicitly separates temporal encoding, multimodal fusion, and scene-level interaction reasoning. Lightweight GRU summaries enable efficient cross-modal attention, while social attention over time--agent tokens captures inter-pedestrian influences at manageable cost. Experiments on JTA, JRDB, and the Pedestrians and Cyclists in Road Traffic dataset show state-of-the-art performance on real-world datasets (JRDB, Urban) and competitive results on JTA. Ablation and qualitative analyses confirm the contribution of each stage and the model's ability to anticipate complex behaviors such as early turning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。