arXiv:2507.08644cs.CV2025-07中稿 · Transactions on In…被引 2

用递归结构高效融合多帧鸟瞰图特征,提升摄像头3D感知性能。

OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception

  • 采用递归结构动态融合多帧鸟瞰图特征,节省内存。
  • 通过运动引导对齐实现跨时序特征精确匹配,提升检测精度。
  • 适合自动驾驶中基于摄像头的实时3D目标检测任务。

基于多视角相机的3D感知可通过透视图到鸟瞰图(BEV)变换获得的BEV特征实现。已有研究证明,结合多帧图像的BEV特征可进一步提升性能。然而,即使补偿了车辆自身运动,当融合大量图像帧时,性能增益仍受限于随时间变化的动态物体引起的BEV特征漂移。本文提出一种新方法OnlineBEV,利用递归结构在时间维度上融合BEV特征,以极低内存开销增加有效融合帧数。为确保跨时序特征的空间对齐,OnlineBEV引入运动引导的鸟瞰图融合网络(MBFNet),从连续的BEV帧中提取运动特征,并据此动态对齐历史与当前特征。同时设计时间一致性学习损失,显式捕捉历史与目标特征间的差异。在nuScenes基准上的实验表明,OnlineBEV在仅使用摄像头的3D目标检测任务中达到63.9%的NDS,超越当前最优方法SOLOFusion,达到领先水平。

原文摘要 · Abstract (English)

Multi-view camera-based 3D perception can be conducted using bird's eye view (BEV) features obtained through perspective view-to-BEV transformations. Several studies have shown that the performance of these 3D perception methods can be further enhanced by combining sequential BEV features obtained from multiple camera frames. However, even after compensating for the ego-motion of an autonomous agent, the performance gain from temporal aggregation is limited when combining a large number of image frames. This limitation arises due to dynamic changes in BEV features over time caused by object motion. In this paper, we introduce a novel temporal 3D perception method called OnlineBEV, which combines BEV features over time using a recurrent structure. This structure increases the effective number of combined features with minimal memory usage. However, it is critical to spatially align the features over time to maintain strong performance. OnlineBEV employs the Motion-guided BEV Fusion Network (MBFNet) to achieve temporal feature alignment. MBFNet extracts motion features from consecutive BEV frames and dynamically aligns historical BEV features with current ones using these motion features. To enforce temporal feature alignment explicitly, we use Temporal Consistency Learning Loss, which captures discrepancies between historical and target BEV features. Experiments conducted on the nuScenes benchmark demonstrate that OnlineBEV achieves significant performance gains over the current best method, SOLOFusion. OnlineBEV achieves 63.9% NDS on the nuScenes test set, recording state-of-the-art performance in the camera-only 3D object detection task.

3D感知鸟瞰图递归融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。