构建首个车路协同3D语义占位预测合成基准,提升自动驾驶感知完整性。
A Synthetic Benchmark for Collaborative 3D Semantic Occupancy Prediction in V2X-Enabled Autonomous Driving
- 在CARLA中设计高分辨率语义体素传感器,生成密集标注数据。
- 提出空间对齐与注意力聚合的多智能体特征融合方法,性能随范围扩大持续提升。
- 提供多尺度评估基准,适合研究协同感知与扩展感知范围的团队使用。
3D语义占位预测是自动驾驶中新兴的感知范式,以体素级表示几何细节与语义类别。然而,单车设置下受限于遮挡、传感器范围和视角狭窄,其效果受制。协作感知通过交换互补信息,可提升预测的完整性和准确性。但相关研究受限于缺乏专用数据集。为此,我们在CARLA中设计高分辨率语义体素传感器,生成密集且全面的标注。进一步开发基线模型,通过空间对齐与注意力聚合实现跨智能体特征融合。同时建立不同预测范围的基准,系统评估空间范围对协作预测的影响。实验表明,基线模型表现优异,且随着预测范围扩大,增益持续增加。代码已开源。
原文摘要 · Abstract (English)
3D semantic occupancy prediction is an emerging perception paradigm in autonomous driving, providing a voxel-level representation of both geometric details and semantic categories. However, its effectiveness is inherently constrained in single-vehicle setups by occlusions, restricted sensor range, and narrow viewpoints. To address these limitations, collaborative perception enables the exchange of complementary information, thereby enhancing the completeness and accuracy of predictions. Despite its potential, research on collaborative 3D semantic occupancy prediction is hindered by the lack of dedicated datasets. To bridge this gap, we design a high-resolution semantic voxel sensor in CARLA to produce dense and comprehensive annotations. We further develop a baseline model that performs inter-agent feature fusion via spatial alignment and attention aggregation. In addition, we establish benchmarks with varying prediction ranges designed to systematically assess the impact of spatial extent on collaborative prediction. Experimental results demonstrate the superior performance of our baseline, with increasing gains observed as range expands. Our code is available at https://github.com/tlab-wide/Co3SOP}{https://github.com/tlab-wide/Co3SOP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。