用跨模态提示增强天气图像修复,提升复杂场景下的还原效果。
Degradation-Aware Prompt Learning with Cross-Modal Compensation for Adverse Weather Removal

- 通过文本与视觉融合生成退化感知提示,指导修复过程。
- 在多个数据集上优于现有方法,结构细节还原更精准。
- 适合需要高保真图像恢复的自动驾驶、遥感领域应用。
恶劣天气引发多样且复杂的图像退化,严重影响计算机视觉系统的可靠性。现有全统一修复模型虽尝试在单一框架内处理多种退化类型,但常缺乏对退化特征的空间与语义建模,限制了其在不同天气条件下的适应性。为此,我们提出一种退化感知跨模态提示补偿网络(DCMPC-Net),利用预训练视觉-语言模型中的跨模态退化线索,在统一主干网络中条件化修复特征。DCMPC-Net主要包括跨模态提示生成器(CMPG)、提示引导注意力对齐模块(PGAAM)和双特征补偿模块(DFCM)。CMPG将文本嵌入与视觉特征融合,生成编码退化相关语义与上下文线索的退化感知提示。这些提示通过PGAAM注入解码器,自适应对齐语义信息与退化区域,促进上下文感知修复。为增强结构保真度,引入DFCM以分离退化伪影与场景结构,从而提升细纹理与细节内容的重建质量。通过整合跨模态语义引导、空间对齐与结构增强,DCMPC-Net在多样化天气条件下实现鲁棒且感知一致的修复。大量实验表明,该方法在任务特定与统一设置下均优于当前最优方法,显著提升准确率与视觉保真度。
原文摘要 · Abstract (English)
Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one restoration models attempt to address multiple degradation types within a unified framework, but often lack explicit spatial and semantic modeling of degradation characteristics, limiting their adaptability to diverse weather conditions. To address this limitation, we propose a Degradation-Aware Cross-Modal Prompt Compensation Network (DCMPC-Net) that leverages cross-modal degradation cues from a pretrained vision-language model to condition restoration features within a unified backbone. Specifically, our DCMPC-Net mainly consists of the Cross-Modal Prompt Generator (CMPG), Prompt-Guided Attention Alignment Module (PGAAM), and Dual Feature Compensation Module (DFCM). The CMPG integrates textual embeddings with visual features to produce degradation-aware prompts that encode degradation-related semantic and contextual cues. These prompts are injected into the decoder via a PGAAM, which adaptively aligns semantic information with degraded regions to facilitate context-aware restoration. To further enhance structural fidelity, DFCM is introduced that disentangles degradation artifacts from scene structures, thereby improving the reconstruction of fine textures and detailed content. By integrating cross-modal semantic guidance with spatial alignment and structural enhancement, DCMPC-Net achieves robust and perceptually consistent restoration across diverse weather conditions. Extensive experiments show that DCMPC-Net outperforms state-of-the-art methods in both task-specific and unified settings, achieving superior accuracy and visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。