arXiv:2507.03893cs.CVcs.AI2025-07中稿 · IEEE Transactions …被引 1

融合可见光与近红外图像,提升远距离雾霾去除效果。

Hierarchical Semantic-Visual Fusion of Visible and Near-infrared Images for Long-range Haze Removal

  • 分层融合可见光与近红外特征,兼顾语义一致性和结构细节。
  • 在真实远距离雾霾场景下,显著减少残留雾霾并恢复清晰纹理。
  • 新构建带语义标注的像素级可见-红外雾霾数据集,便于评测。

尽管图像去雾在过去十年中取得显著进展,但多数研究集中在短距离场景,远距离去雾仍缺乏探索。随着距离增加,散射增强导致严重雾霾和信号损失,仅靠可见光图像难以恢复远距离细节。近红外波段具有更强的雾穿透能力,通过多模态融合可提供关键互补信息。然而,现有方法多关注内容整合,常忽略可见光图像中的雾霾特征,导致结果残留雾霾。本文提出分层语义-视觉融合(HSVF)框架,认为可见光与近红外模态不仅提供互补的低层视觉特征,还共享高层语义一致性。为此,构建语义流以重建无雾场景,视觉流则从近红外图像中恢复丢失的结构细节。语义流首先通过对齐模态不变的内在表示,获得抗雾霾的语义预测;共享语义作为强先验,用于恢复严重雾霾下的高对比度远距离场景。同时,视觉流融合可见光与近红外的互补线索,恢复丰富纹理。双流协作使结果兼具高对比度与细腻纹理。此外,我们构建了一个新型像素级可见-红外雾霾数据集,含语义标签,用于基准评测。大量实验证明,本方法在真实远距离去雾任务上优于当前最优技术。

原文摘要 · Abstract (English)

While image dehazing has advanced substantially in the past decade, most efforts have focused on short-range scenarios, leaving long-range haze removal under-explored. As distance increases, intensified scattering leads to severe haze and signal loss, making it impractical to recover distant details solely from visible images. Near-infrared, with superior fog penetration, offers critical complementary cues through multimodal fusion. However, existing methods focus on content integration while often neglecting haze embedded in visible images, leading to results with residual haze. In this work, we argue that the infrared and visible modalities not only provide complementary low-level visual features, but also share high-level semantic consistency. Motivated by this, we propose a Hierarchical Semantic-Visual Fusion (HSVF) framework, comprising a semantic stream to reconstruct haze-free scenes and a visual stream to incorporate structural details from the near-infrared modality. The semantic stream first acquires haze-robust semantic prediction by aligning modality-invariant intrinsic representations. Then the shared semantics act as strong priors to restore clear and high-contrast distant scenes under severe haze degradation. In parallel, the visual stream focuses on recovering lost structural details from near-infrared by fusing complementary cues from both visible and near-infrared images. Through the cooperation of dual streams, HSVF produces results that exhibit both high-contrast scenes and rich texture details. Moreover, we introduce a novel pixel-aligned visible-infrared haze dataset with semantic labels to facilitate benchmarking. Extensive experiments demonstrate the superiority of our method over state-of-the-art approaches in real-world long-range haze removal.

去雾多模态融合远距离近红外

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。