arXiv:2511.20020cs.CV2025-11被引 3

用多模态注意力机制预测行人过街意图,提升自动驾驶安全性。

ACIT: Attention-Guided Cross-Modal Interaction Transformer for Pedestrian Crossing Intention Prediction

  • 设计双路径注意力机制融合视觉与运动模态数据
  • 在JAADbeh和JAADall数据集上分别达到70%和89%准确率
  • 适合关注行人行为预测与多模态融合的自动驾驶研究者

预测行人过街意图对自动驾驶车辆避免碰撞至关重要。然而,有效提取并整合不同数据类型的互补信息仍是主要挑战。本文提出一种注意力引导的跨模态交互Transformer(ACIT),利用六种视觉与运动模态,分为三组交互对:(1) 全局语义图与全局光流,(2) 局部RGB图像与局部光流,(3) 自车速度与行人边界框。每组视觉交互对中,双路径注意力机制通过模内自注意力增强主模态显著区域,并通过光流引导注意力实现与辅助模态(光流)的深度交互。运动交互对中,采用跨模态注意力建模动态关系,有效提取互补运动特征。此外,多模态特征融合模块在每个时间步促进跨模态交互,Transformer-based时序聚合模块捕捉序列依赖。实验表明,ACIT优于现有方法,在JAADbeh和JAADall数据集上分别达到70%和89%准确率。大量消融实验证明各模块贡献。

原文摘要 · Abstract (English)

Predicting pedestrian crossing intention is crucial for autonomous vehicles to prevent pedestrian-related collisions. However, effectively extracting and integrating complementary cues from different types of data remains one of the major challenges. This paper proposes an attention-guided cross-modal interaction Transformer (ACIT) for pedestrian crossing intention prediction. ACIT leverages six visual and motion modalities, which are grouped into three interaction pairs: (1) Global semantic map and global optical flow, (2) Local RGB image and local optical flow, and (3) Ego-vehicle speed and pedestrian's bounding box. Within each visual interaction pair, a dual-path attention mechanism enhances salient regions within the primary modality through intra-modal self-attention and facilitates deep interactions with the auxiliary modality (i.e., optical flow) via optical flow-guided attention. Within the motion interaction pair, cross-modal attention is employed to model the cross-modal dynamics, enabling the effective extraction of complementary motion features. Beyond pairwise interactions, a multi-modal feature fusion module further facilitates cross-modal interactions at each time step. Furthermore, a Transformer-based temporal feature aggregation module is introduced to capture sequential dependencies. Experimental results demonstrate that ACIT outperforms state-of-the-art methods, achieving accuracy rates of 70% and 89% on the JAADbeh and JAADall datasets, respectively. Extensive ablation studies are further conducted to investigate the contribution of different modules of ACIT.

行人预测多模态融合Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。