arXiv:2409.11706cs.CV2024-09被引 6

首个面向路侧多相机鸟瞰感知的稠密方法,解决姿态多样、视角稀疏等难题。

RopeBEV: A Multi-Camera Roadside Perception Network in Bird's-Eye-View

  • 通过鸟瞰图增强缓解不同摄像头姿态带来的训练不平衡问题。
  • 在真实高速数据集RoScenes上排名第一,覆盖600个摄像头和50多个路口。
  • 适合大规模路侧智能交通系统部署,尤其支持动态摄像头数量变化。

车载多摄像头鸟瞰(BEV)感知方法已在自动驾驶中广泛应用,但路侧场景因摄像头布局与车载场景差异显著,尚无成熟的多相机BEV解决方案。本文系统分析了路侧场景下多摄像头BEV感知的关键挑战:摄像头位姿多样性、摄像头数量不确定性、感知区域稀疏性以及方向角模糊性。为此,提出RopeBEV,首个面向路侧的稠密多相机BEV方法。引入BEV增强以缓解由多样化摄像头位姿导致的训练不平衡;通过CamMask与ROIMask分别支持可变摄像头数量和稀疏感知区域;利用摄像头旋转嵌入解决方向模糊问题。该方法在真实高速公路数据集RoScenes上取得第一名,并在包含超过50个路口和600个摄像头的私有城市数据集上验证了实际应用价值。

原文摘要 · Abstract (English)

Multi-camera perception methods in Bird's-Eye-View (BEV) have gained wide application in autonomous driving. However, due to the differences between roadside and vehicle-side scenarios, there currently lacks a multi-camera BEV solution in roadside. This paper systematically analyzes the key challenges in multi-camera BEV perception for roadside scenarios compared to vehicle-side. These challenges include the diversity in camera poses, the uncertainty in Camera numbers, the sparsity in perception regions, and the ambiguity in orientation angles. In response, we introduce RopeBEV, the first dense multi-camera BEV approach. RopeBEV introduces BEV augmentation to address the training balance issues caused by diverse camera poses. By incorporating CamMask and ROIMask (Region of Interest Mask), it supports variable camera numbers and sparse perception, respectively. Finally, camera rotation embedding is utilized to resolve orientation ambiguity. Our method ranks 1st on the real-world highway dataset RoScenes and demonstrates its practical value on a private urban dataset that covers more than 50 intersections and 600 cameras.

路侧感知鸟瞰图多相机自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。