arXiv:2511.20008cs.CVcs.AI2025-11

用多模态融合预测行人过街意图,提升自动驾驶安全性。

Pedestrian Crossing Intention Prediction Using Multimodal Fusion Network

  • 结合视觉与运动特征,通过Transformer提取多模态信息。
  • 在JAAD数据集上优于基线方法,显著提升预测准确率。
  • 适合自动驾驶场景中行人行为理解研究者参考。

行人过街意图预测对自动驾驶车辆在城市环境中的部署至关重要。理想的预测能为车辆提供关键环境线索,从而降低与行人相关的碰撞风险。然而,由于行人行为多样且依赖多种上下文因素,该任务具有挑战性。本文提出一种多模态融合网络,利用来自视觉和运动分支的七种模态特征,有效提取并整合不同模态间的互补线索。具体而言,使用多个基于Transformer的提取模块从原始输入中获取运动与视觉特征;深度引导注意力模块通过全面的空间特征交互,利用深度信息引导另一模态的关注区域。为应对不同模态和帧的重要性差异,设计了模态注意力与时间注意力机制,以选择性强调有用模态并有效捕捉时间依赖关系。在JAAD数据集上的大量实验验证了所提网络的有效性,性能优于基线方法。

原文摘要 · Abstract (English)

Pedestrian crossing intention prediction is essential for the deployment of autonomous vehicles (AVs) in urban environments. Ideal prediction provides AVs with critical environmental cues, thereby reducing the risk of pedestrian-related collisions. However, the prediction task is challenging due to the diverse nature of pedestrian behavior and its dependence on multiple contextual factors. This paper proposes a multimodal fusion network that leverages seven modality features from both visual and motion branches, aiming to effectively extract and integrate complementary cues across different modalities. Specifically, motion and visual features are extracted from the raw inputs using multiple Transformer-based extraction modules. Depth-guided attention module leverages depth information to guide attention towards salient regions in another modality through comprehensive spatial feature interactions. To account for the varying importance of different modalities and frames, modality attention and temporal attention are designed to selectively emphasize informative modalities and effectively capture temporal dependencies. Extensive experiments on the JAAD dataset validate the effectiveness of the proposed network, achieving superior performance compared to the baseline methods.

行人预测多模态融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。