打造最大规模的车路协同感知数据集,助力自动驾驶多智能体协作
SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception
- 基于CARLA模拟器自动生成多样化驾驶场景,同步采集多模态数据
- 数据集含258个场景、超2700万标注框,规模超现有数据集一个数量级
- 提供新模型CoBEVFusion,提升多车协同感知精度,适合研究者使用
通过车路协同(V2X)通信可突破单个自动驾驶车辆在遮挡和传感器范围上的物理限制。然而,依赖统一空间表示(如鸟瞰图)的鲁棒性算法开发受限于缺乏大规模、多模态、多任务数据集。真实世界多智能体同步数据的采集与标注成本极高,导致现有数据集规模与范围均有限。为此,我们提出SimBEV2X——一个基于CARLA模拟器的先进合成数据生成工具,能自动生成随机驾驶场景,采集多模态传感器数据及多种地面真值信息,包括3D边界框(带唯一轨迹ID)、高清地图、鸟瞰图分割图、语义占用体素网格(来自车辆与路侧单元RSU)。我们构建了当前最大的V2X感知数据集——SimBEV2X,包含258个场景,每场景最多8辆联网车与4个RSU,覆盖多样道路网络。该数据集共含102,200帧、588,520个激光点云、超300万张图像、超过2700万条边界框及全面标注。此外,我们以CoopDet3D为基线,提出CoBEVFusion——一种结合融合轴向注意力(FAX)的新型架构,实现上下文感知的多智能体特征聚合,性能显著提升。SimBEV2X数据集、工具及代码已开源。
原文摘要 · Abstract (English)
Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms, particularly those relying on unified spatial representations like bird's-eye view (BEV) representation, is hampered by the lack of large-scale, multi-modal, multi-task datasets. Moreover, collecting and annotating a large set of synchronized, real-world multi-agent data is prohibitively expensive. This has resulted in a landscape where existing V2X datasets are notably limited in both size and scope. To overcome this, we introduce SimBEV2X, an advanced synthetic data generation tool built on the CARLA simulator. SimBEV2X automatically creates randomized driving scenarios to collect multi-modal sensor data alongside various types of ground truth including 3D bounding boxes with unique track IDs, HD map information, BEV segmentation maps, and semantic occupancy voxel grids from both vehicles and RSUs. We also present the SimBEV2X dataset, the largest V2X perception dataset to date. The dataset comprises 258 scenes, each involving up to 8 connected vehicles and up to 4 RSUs across a variety of road networks. The SimBEV2X dataset is an order of magnitude larger than existing V2X datasets and contains 102,200 frames, 588,520 lidar point clouds, more than 3 million images, over 27 million bounding boxes, and a comprehensive set of other annotations. Finally, we establish a strong baseline on the SimBEV2X dataset using CoopDet3D and propose CoBEVFusion, a novel architecture that combines CoopDet3D with fused axial attention (FAX) for context-aware multi-agent feature aggregation, resulting in superior performance. SimBEV2X, the SimBEV2X dataset, and CoBEVFusion are available at https://simbev2x.org and https://github.com/GoodarzMehr/SimBEV2X.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。