构建了分级雾霾下的多模态自动驾驶数据集,助力感知模型在恶劣天气下更可靠。
FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog

- 用物理模型模拟三种能见度的雾霾,同步生成相机、激光雷达等多模态数据
- 133万帧标注数据中车辆检测精度达95.1%,距离40米内召回率超99%
- 揭示了训练时混合不同密度雾可提升3D检测性能,且无需增加数据量
恶劣天气下的感知仍是自动驾驶可靠性的关键瓶颈,现有基准缺乏系统化的多模态对齐。真实世界数据集受采集条件不可控和单一、未校准的天气状态限制,而合成数据要么仅针对摄像头复原,要么缺少成对的清晰与雾霾场景,无法评估“去雾后检测”流程。我们提出FogDrive,一个严格校准的多模态自动驾驶数据集,融合数据工程与鲁棒机器学习。基于CARLA仿真器构建,包含660个场景(约13.3万帧全标注图像,昼夜各半),四台同步相机(RGB、深度、语义分割)、激光雷达及语义激光雷达对、前向雷达。使用Koschmieder模型(相机)和Beer-Lambert定律(激光雷达)分别建模物理一致的雾霾,设定三个校准能见度等级(160米、100米、50米)。每个场景提供四种匹配版本(清晰+三档雾霾),并附带跨模态2D/3D边界框。基于语义分割的质量审计显示,在8000张图像上车辆检测精度为95.1%,40米内召回率超过99%。我们以TransFusion、BEVFusion、YOLOv8-m等先进架构建立基准,涵盖3D多模态融合与2D图像复原两种范式。结果揭示:训练时混合多密度雾可提升3D边界框几何精度,且无需额外数据成本;而在2D流程中,图像质量指标(如PSNR、SSIM)难以预测下游检测表现。FogDrive将与数据生成框架一并开源,推动多模态鲁棒研究发展。
原文摘要 · Abstract (English)
Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled collection and single-level, uncalibrated conditions, while synthetic alternatives either target camera-only restoration or lack the paired clean-and-foggy structure needed to benchmark "defog-then-detect" pipelines. We present FogDrive, a rigorously calibrated, multi-modal autonomous-driving dataset bridging data-centric engineering and robust machine learning. Built with the CARLA simulator, FogDrive contains 660 scenes (~133k fully annotated frames, 50:50 day/night) across four synchronized cameras (RGB, depth, semantic segmentation), a LiDAR and semantic-LiDAR pair, and front radar. Physically consistent fog is modeled independently on camera channels (Koschmieder model) and LiDAR channels (Beer-Lambert law) at three calibrated visibility densities (160m, 100m, 50m). Every scene ships in four matched variants (clean plus three graded fog levels) with cross-calibrated 2D and 3D bounding boxes. A semantic-segmentation-based quality audit over 8k images validates annotations at 95.1% precision and over 99% recall for vehicles within 40m. We establish baseline benchmarks with state-of-the-art architectures (TransFusion, BEVFusion, YOLOv8-m) across two paradigms: 3D multi-modal fusion and 2D image restoration. These yield critical data-centric insights: mixing multi-density fog during training tightens 3D bounding-box geometry without added data-scaling cost, while in 2D pipelines image-quality metrics (PSNR, SSIM) prove poor predictors of downstream detection performance. FogDrive will be fully open-sourced alongside our data-generation framework to accelerate robust, multi-modal research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。