用传感器姿态引导多模态数据对齐,减少标注依赖。
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
- 以传感器姿态为监督信号,对齐相机与激光雷达的BEV特征。
- 仅需少量标注数据,在BEV分割任务上超越全监督方法。
- 适合数据稀缺场景,如非城市、野外和室内环境应用。
在自动驾驶与移动机器人领域,构建鸟瞰图(BEV)表示的方法正从传统方式转向基于Transformer的多模态融合,主要将激光雷达与摄像头的数据融合为二维地面平面表征。然而,这些学习型方法通常严重依赖大量标注数据,尤其在缺乏大规模数据集的多样化或非城市环境中面临挑战。本文提出BEVPose框架,利用传感器姿态作为指导性监督信号,融合相机与激光雷达的BEV表示。通过姿态信息对齐多模态输入,促进学习隐式BEV嵌入,同时捕捉环境的几何与语义特征。预训练方法在BEV地图分割任务中表现优异,优于现有全监督先进方法,且仅需极少标注数据。该工作不仅提升了BEV表示学习的数据效率,也拓展了其在非道路与室内等场景的应用潜力。
原文摘要 · Abstract (English)
In the field of autonomous driving and mobile robotics, there has been a significant shift in the methods used to create Bird's Eye View (BEV) representations. This shift is characterised by using transformers and learning to fuse measurements from disparate vision sensors, mainly lidar and cameras, into a 2D planar ground-based representation. However, these learning-based methods for creating such maps often rely heavily on extensive annotated data, presenting notable challenges, particularly in diverse or non-urban environments where large-scale datasets are scarce. In this work, we present BEVPose, a framework that integrates BEV representations from camera and lidar data, using sensor pose as a guiding supervisory signal. This method notably reduces the dependence on costly annotated data. By leveraging pose information, we align and fuse multi-modal sensory inputs, facilitating the learning of latent BEV embeddings that capture both geometric and semantic aspects of the environment. Our pretraining approach demonstrates promising performance in BEV map segmentation tasks, outperforming fully-supervised state-of-the-art methods, while necessitating only a minimal amount of annotated data. This development not only confronts the challenge of data efficiency in BEV representation learning but also broadens the potential for such techniques in a variety of domains, including off-road and indoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。