融合4D雷达与相机数据,提升复杂环境下3D语义占位预测鲁棒性。
4DRC-OCC: Robust Semantic Occupancy Prediction Through Fusion of 4D Radar and Camera
- 用4D雷达和相机互补信息实现多模态融合
- 在恶劣天气下仍保持高精度占位预测
- 自动生成标注数据集,降低人工成本
自动驾驶需在各种环境条件下具备鲁棒感知能力,但现有3D语义占位预测在恶劣天气和光照条件下仍具挑战。本文首次研究结合4D雷达与相机数据进行3D语义占位预测。融合策略利用两者互补优势:4D雷达在复杂条件下提供可靠的测距、速度与角度信息,相机则提供丰富的语义与纹理细节。进一步表明,通过相机像素的深度线索可将2D图像升维至3D,显著提升场景重建精度。此外,我们构建了一个全自动标注的数据集,大幅减少对昂贵人工标注的依赖。实验验证了4D雷达在多样化场景中的鲁棒性,凸显其推动自动驾驶感知技术发展的潜力。
原文摘要 · Abstract (English)
Autonomous driving requires robust perception across diverse environmental conditions, yet 3D semantic occupancy prediction remains challenging under adverse weather and lighting. In this work, we present the first study combining 4D radar and camera data for 3D semantic occupancy prediction. Our fusion leverages the complementary strengths of both modalities: 4D radar provides reliable range, velocity, and angle measurements in challenging conditions, while cameras contribute rich semantic and texture information. We further show that integrating depth cues from camera pixels enables lifting 2D images to 3D, improving scene reconstruction accuracy. Additionally, we introduce a fully automatically labeled dataset for training semantic occupancy models, substantially reducing reliance on costly manual annotation. Experiments demonstrate the robustness of 4D radar across diverse scenarios, highlighting its potential to advance autonomous vehicle perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。