用实验室数据训练,再合成雾化数据微调,实现跨场景清晰成像。
From Fog Chamber to Aircraft Window: Pixel-Registered Imaging and Synthetic Fine-Tuning Enable Cross-Domain Defogging

- 在固定路径的雾室中采集像素对齐的雾图与清晰图,支持精确重建。
- 合成雾数据微调后,在机窗视频上图像质量提升,且时序稳定。
- 适合需要跨域泛化的低光照/雾天视觉任务研究者使用。
一种深度去雾流水线在受控实验室雾环境中预训练,并通过域随机合成雾数据进行微调,应用于无目标域训练的室外场景,可泛化至从舱内自由流动雾到飞行中透过飞机舷窗拍摄的手机视频这一系列分布外设置,涵盖全新传感器、场景与光学路径。该方法直接解决了真实世界双目去雾中的开放性迁移限制。两个设计支撑迁移能力:第一,单相机通过固定114毫米散射路径拍摄平板显示的雾室图像,生成5,495对像素对齐的雾图/清晰图,精确配准支持基于拉普拉斯比的恢复质量预测(斯皮尔曼ρ=0.632,优于单图代理的0.399),并支持像素级精确的L1重建训练,避免对抗幻觉;第二,雾室初始模型在Mapillary Vistas图像块上叠加实时随机合成雾(覆盖广泛强度、空间变化、光照和噪声)进行微调。在552张独立测试集上,30种修复骨干网络对比显示NAFNet表现最佳(24.33 dB / 0.7912 SSIM),一个紧凑替代方案仅差1.29 dB,参数量仅为3%;ResNet-50分类器验证了恢复结果保留语义内容而非仅像素结构。在无配对的机窗视频上,NIQE从均值6.22降至4.97,输出时序稳定。同一骨干在有配对监督下,也在非重叠的O-HAZE/NH-HAZE分割上达到20.71 dB / 0.683 SSIM(作为迁移能力检验而非排名竞争)。
原文摘要 · Abstract (English)
A deep defogging pipeline pretrained on controlled laboratory fog and fine-tuned with domain-randomized synthetic fog applied to clear outdoor scenes generalizes across a graded sequence of out-of-distribution settings with no target-domain training, from chamber-free free-flowing fog to iPhone video recorded through an aircraft cabin window in flight, an entirely unseen sensor, scene, and optical path. This directly addresses an open transfer limitation reported for real-world binocular defogging. Two design choices support the transfer. First, a single-camera fog imager photographs a flat-panel display through an artificial-fog enclosure with a fixed 114~mm scattering path, producing 5{,}495 pixel-aligned foggy/clear pairs. Exact registration permits a paired Laplacian ratio that predicts per-image restoration quality far better than single-image proxies (Spearman $ρ= 0.632$ versus $0.399$) and supports pixel-exact $L_1$ reconstruction training that avoids adversarial hallucination. Second, the fog-chamber checkpoint is fine-tuned on Mapillary Vistas crops overlaid with on-the-fly randomized synthetic fog spanning a broad range of strengths, spatial variations, airlights, and noise conditions. On a 552-image held-out split, a uniform comparison of 30 restoration backbones places NAFNet at the top (24.33~dB~/~0.7912~SSIM), with a compact alternative within 1.29~dB at 3\% of the parameter count, and a ResNet-50 classifier confirms that the restoration preserves semantic content rather than only pixel-level structure. On unpaired aircraft-window video, NIQE decreases from a mean of 6.22 to 4.97 after fine-tuning, with temporally stable output across full-motion sequences. The same backbone, under paired supervision, also reaches 20.71~dB~/~0.683~SSIM on a non-overlapping O-HAZE/NH-HAZE split (a transferability check rather than a competitive ranking).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。