对比多种去雾方法在真实与合成数据上的表现,发现并非所有去雾都提升自动驾驶感知能力。
From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- 用真实和合成雾天数据评估传统滤波、深度网络及视觉语言模型的去雾效果
- 发现去雾后检测与分割性能提升仅在特定条件下成立,且模型组合效果不一
- 首次引入人类与AI评委的定性评估,验证其与下游任务指标的一致性
自动驾驶感知系统在雾天环境下尤为脆弱,因光线散射导致对比度下降、细节模糊,影响安全运行。尽管已有大量去雾方法,从手工滤波器到学习型恢复模型,但图像质量提升并不总能转化为下游检测与分割性能的改善。以往评估多依赖合成数据,引发真实场景泛化能力的担忧。本文系统性地评估了涵盖经典去雾滤波器、现代去雾网络、滤波器与模型串联组合、以及直接作用于雾图的提示驱动视觉语言模型在内的完整去雾流程。为弥合模拟与真实环境差距,我们在合成的Foggy Cityscapes数据集与真实世界的Adverse Conditions Dataset with Correspondences(ACDC)上进行评估。通过分析合成雾与真实天气下的泛化能力,以目标检测mAP和语义分割全景质量作为下游任务指标,研究去雾的有效性、模型组合的影响及视觉语言模型的表现。我们还报告了基于定性评分的人类与视觉语言模型评委结果,并分析其与任务指标的一致性。这些成果建立了一个透明、任务导向的去雾基准,明确了预处理在恶劣天气下提升自动驾驶感知的适用条件。
原文摘要 · Abstract (English)
Autonomous driving perception systems are particularly vulnerable in foggy conditions, where light scattering reduces contrast and obscures fine details critical for safe operation. While numerous defogging methods exist, from handcrafted filters to learned restoration models, improvements in image fidelity do not consistently translate into better downstream detection and segmentation. Moreover, prior evaluations often rely on synthetic data, raising concerns about real-world transferability. We present a structured empirical study that benchmarks a comprehensive set of defogging pipelines, including classical dehazing filters, modern defogging networks, chained variants combining filters and models, and prompt-driven visual language image editing models applied directly to foggy images. To bridge the gap between simulated and physical environments, we evaluate these pipelines on both the synthetic Foggy Cityscapes dataset and the real-world Adverse Conditions Dataset with Correspondences (ACDC). We examine generalization by evaluating performance on synthetic fog and real-world conditions, assessing both image quality and downstream perception in terms of object detection mean average precision and segmentation panoptic quality. Our analysis identifies when defogging is effective, the impact of combining models, and how visual language models compare to traditional approaches. We additionally report qualitative rubric-based evaluations from both human and visual language model judges and analyze their alignment with downstream task metrics. Together, these results establish a transparent, task-oriented benchmark for defogging methods and identify the conditions under which pre-processing meaningfully improves autonomous perception in adverse weather. Project page: https://aradfir.github.io/filters-to-vlms-defogging-page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。