用CLIP筛选外部数据,提升夜间去雾的训练稳定性与效果
CLIP-Guided Data Augmentation for Night-Time Image Dehazing

- 用预训练CLIP筛选相似外部图像,构建更贴近目标域的训练数据
- 分阶段训练NAFNet,先适配夜间域再扩展至复杂退化模式
- 推理时结合自集成与加权融合,提升输出一致性与视觉质量
夜间图像去雾面临比白天更复杂的退化问题,因雾霾散射与低光照、非均匀照明及强光干扰耦合。在监督有限的情况下,这种复杂性加剧领域偏移和训练不稳定性,因目标域样本稀缺,而盲目引入外部数据又会因分布差异削弱适应能力。本文针对NTIRE 2026夜间图像去雾挑战提出统一框架,包含域对齐数据构建、分阶段训练与推理时增强。具体地,预训练的CLIP视觉编码器通过相似性筛选候选外部样本,构建更接近目标域的训练数据;随后使用NAFNet进行两阶段训练:先适应目标域,再扩展至更广泛的退化模式;推理时结合TLC、x8自集成与加权快照融合,提升输出稳定性。该框架不依赖复杂网络重设计,提供一种实用且高效的夜间去雾方案。
原文摘要 · Abstract (English)
Nighttime image dehazing faces a more complex degradation pattern than its daytime counterpart, as haze scattering couples with low illumination, non-uniform lighting, and strong light interference. Under limited supervision, this complexity aggravates domain drift and training instability, since target-domain samples are scarce while naively introducing external data may weaken adaptation due to distribution mismatch. This paper presents our solution to the NTIRE 2026 Night Time Image Dehazing Challenge, built as a unified framework that integrates domain-aligned data construction, stage-wise training, and inference-time enhancement. Specifically, a pre-trained CLIP visual encoder screens candidate external samples by similarity to construct training data closer to the target domain. NAFNet is then trained in two stages, first adapting to the target domain and then expanding to broader degradation patterns. At inference time, TLC, x8 self-ensemble, and weighted snapshot fusion are combined to improve output stability. Rather than relying on complex network redesign, the proposed framework offers a practical and effective pipeline for nighttime image dehazing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。