arXiv:2604.02930cs.CV2026-04被引 1

用注意力机制提升自动驾驶鸟瞰图预测精度

BEVPredFormer: Spatio-temporal Attention for BEV Instance Prediction in Autonomous Driving

  • 采用分时空间注意力与门控变压器结构,捕捉动态场景细粒度运动
  • 在nuScenes数据集上达到或超越当前最优性能,支持实时推理
  • 适合需要高精度行为预测的自动驾驶感知系统

自动驾驶系统需准确感知动态场景演化,以检测、追踪并预测周围障碍物行为。传统模块化感知流程易产生累积误差和延迟。实例预测模型通过融合多传感器信息,在鸟瞰图(BEV)中统一完成当前及未来帧的分割与运动估计。然而,动态驾驶环境中密集的空间-时间信息带来挑战,要求模型在不牺牲实时性的前提下,捕捉精细运动模式与长程依赖。本文提出BEVPredFormer,一种纯摄像头输入的新型BEV实例预测架构,通过基于注意力的时序处理增强时空理解,并采用注意力驱动的3D投影。其采用无循环设计,结合门控变压器层、分时空间注意力机制与多尺度头任务,还引入差异引导特征提取模块以强化时序表征。大量消融实验验证各组件有效性。在nuScenes数据集上的评估表明,BEVPredFormer性能达到或超过现有最先进方法,展现出在鲁棒高效自动驾驶感知中的潜力。

原文摘要 · Abstract (English)

A robust awareness of how dynamic scenes evolve is essential for Autonomous Driving systems, as they must accurately detect, track, and predict the behaviour of surrounding obstacles. Traditional perception pipelines that rely on modular architectures tend to suffer from cumulative errors and latency. Instance Prediction models provide a unified solution, performing Bird's-Eye-View segmentation and motion estimation across current and future frames using information directly obtained from different sensors. However, a key challenge in these models lies in the effective processing of the dense spatial and temporal information inherent in dynamic driving environments. This level of complexity demands architectures capable of capturing fine-grained motion patterns and long-range dependencies without compromising real-time performance. We introduce BEVPredFormer, a novel camera-only architecture for BEV instance prediction that uses attention-based temporal processing to improve temporal and spatial comprehension of the scene and relies on an attention-based 3D projection of the camera information. BEVPredFormer employs a recurrent-free design that incorporates gated transformer layers, divided spatio-temporal attention mechanisms, and multi-scale head tasks. Additionally, we incorporate a difference-guided feature extraction module that enhances temporal representations. Extensive ablation studies validate the effectiveness of each architectural component. When evaluated on the nuScenes dataset, BEVPredFormer was on par or surpassed State-Of-The-Art methods, highlighting its potential for robust and efficient Autonomous Driving perception.

自动驾驶目标预测注意力机制鸟瞰图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。