仅用一张全景图实现鸟瞰语义地图,简化系统设计并提升精度。
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
- 用单张全景图替代多摄像头输入,避免标定同步难题。
- 提出MVT模块,有效处理全景图空间畸变,提升特征转换精度。
- 在nuScenes-360和DeepAccident-360上分别达到51.1%和36.1% mIoU,性能领先。
在自动驾驶领域,鸟瞰视图(BEV)感知因提供比针孔前视图和全景图更全面的信息而受到广泛关注。传统BEV方法依赖多个窄视角摄像头及复杂的姿态估计,常面临标定与同步问题。为突破上述挑战,本文提出OneBEV,一种仅需单张全景图像作为输入的新型BEV语义映射方法,简化映射流程并降低计算复杂度。专门设计了名为Mamba View Transformation(MVT)的畸变感知模块,用于处理全景图中的空间畸变,将前视特征转换为BEV特征,无需依赖传统注意力机制。此外,本文还构建了两个专用于OneBEV任务的数据集:nuScenes-360和DeepAccident-360。实验结果表明,OneBEV在nuScenes-360和DeepAccident-360上的平均交并比(mIoU)分别达到51.1%和36.1%,表现达到当前最优水平。该工作推动了自动驾驶中BEV语义映射的发展,为更先进可靠的自动驾驶系统铺平道路。
原文摘要 · Abstract (English)
In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared with pinhole front-view images and panoramas. Traditional BEV methods, which rely on multiple narrow-field cameras and complex pose estimations, often face calibration and synchronization issues. To break the wall of the aforementioned challenges, in this work, we introduce OneBEV, a novel BEV semantic mapping approach using merely a single panoramic image as input, simplifying the mapping process and reducing computational complexities. A distortion-aware module termed Mamba View Transformation (MVT) is specifically designed to handle the spatial distortions in panoramas, transforming front-view features into BEV features without leveraging traditional attention mechanisms. Apart from the efficient framework, we contribute two datasets, i.e., nuScenes-360 and DeepAccident-360, tailored for the OneBEV task. Experimental results showcase that OneBEV achieves state-of-the-art performance with 51.1% and 36.1% mIoU on nuScenes-360 and DeepAccident-360, respectively. This work advances BEV semantic mapping in autonomous driving, paving the way for more advanced and reliable autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。