arXiv:2511.13113cs.CV2025-11AAAI被引 1

用语义与结构先验提升去雨图像的细节保真度

Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining

  • 融合CLIP语义与DINOv2结构先验,分层注入互补信息
  • 在Rain200H上提升0.57 dB PSNR,真实雨天场景泛化更强
  • 适合需要高保真去雨的自动驾驶与监控系统

雨水显著降低计算机视觉系统性能,尤其在自动驾驶和视频监控中。现有去雨方法常难以保持语义与空间细节的真实性。为此,我们提出多先验分层Mamba(MPHM)网络用于图像去雨。该架构协同整合宏观语义文本先验(CLIP)实现任务级语义引导,以及微观结构视觉先验(DINOv2)提供场景感知的结构信息。为缓解异构先验间的潜在冲突,设计渐进式先验融合注入(PFI),在解码器不同层级策略性地注入互补线索。同时,将精细的分层Mamba模块(HMM)引入主干网络,其傅里叶增强双路径设计能同步建模全局上下文与恢复局部细节。大量实验表明,MPHM达到领先性能,在Rain200H数据集上获得0.57 dB PSNR提升,并在真实雨天场景中表现出更优泛化能力。

原文摘要 · Abstract (English)

Rain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining methods have made considerable progress, they often struggle with fidelity of semantic and spatial details. To address these limitations, we propose the Multi-Prior Hierarchical Mamba (MPHM) network for image deraining. This novel architecture synergistically integrates macro-semantic textual priors (CLIP) for task-level semantic guidance and micro-structural visual priors (DINOv2) for scene-aware structural information. To alleviate potential conflicts between heterogeneous priors, we devise a progressive Priors Fusion Injection (PFI) that strategically injects complementary cues at different decoder levels. Meanwhile, we equip the backbone network with an elaborate Hierarchical Mamba Module (HMM) to facilitate robust feature representation, featuring a Fourier-enhanced dual-path design that concurrently addresses global context modeling and local detail recovery. Comprehensive experiments demonstrate MPHM's state-of-the-art performance, achieving a 0.57 dB PSNR gain on the Rain200H dataset while delivering superior generalization on real-world rainy scenarios.

图像去雨多先验Mamba结构感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。