针对水下图像增强中背景干扰问题,提出动态重排扫描机制提升关键区域特征表达。
VRS-UIE: Value-Driven Reordering Scanning for Underwater Image Enhancement
- 根据像素价值动态重排扫描顺序,优先处理重要目标区域。
- 在水下图像上平均提升0.89 dB,有效抑制水色偏移并保留结构与色彩。
- 轻量版适配实时应用,兼顾性能与效率,适合部署在边缘设备。
状态空间模型(SSMs)因其线性复杂度和全局感受野,在视觉任务中展现出巨大潜力。然而,在水下图像增强(UIE)任务中,标准的顺序扫描机制受到水下场景独特统计分布特性的挑战:大面积同质但无用的海洋背景会稀释稀疏但有价值的目標特征响应,阻碍有效状态传播,削弱模型对局部语义和全局结构的保持能力。为此,我们提出一种新的价值驱动重排扫描框架——VRS-UIE。其核心是多粒度价值引导学习(MVGL)模块,生成像素级价值图以动态重排SSM扫描序列,优先传递显著特征的长程状态信息。在此基础上,设计了融合优先级驱动全局排序与动态调整局部卷积的Mamba-Conv Mixer(MCM)模块,有效建模大范围海洋背景与高价值语义目标。进一步引入跨特征桥接(CFB)优化多层次特征融合。大量实验表明,所提框架达到新最优性能,平均比WMamba提升0.89 dB,有效抑制水色偏移,同时保持结构与色彩保真度。此外,通过引入高效卷积算子与分辨率缩放策略,构建出轻量高效的VRS-UIE-S方案,适用于实时水下图像增强应用。
原文摘要 · Abstract (English)
State Space Models (SSMs) have emerged as a promising backbone for vision tasks due to their linear complexity and global receptive field. However, in the context of Underwater Image Enhancement (UIE), the standard sequential scanning mechanism is fundamentally challenged by the unique statistical distribution characteristics of underwater scenes. The predominance of large-portion, homogeneous but useless oceanic backgrounds can dilute the feature representation responses of sparse yet valuable targets, thereby impeding effective state propagation and compromising the model's ability to preserve both local semantics and global structure. To address this limitation, we propose a novel Value-Driven Reordering Scanning framework for UIE, termed VRS-UIE. Its core innovation is a Multi-Granularity Value Guidance Learning (MVGL) module that generates a pixel-aligned value map to dynamically reorder the SSM's scanning sequence. This prioritizes informative regions to facilitate the long-range state propagation of salient features. Building upon the MVGL, we design a Mamba-Conv Mixer (MCM) block that synergistically integrates priority-driven global sequencing with dynamically adjusted local convolutions, thereby effectively modeling both large-portion oceanic backgrounds and high-value semantic targets. A Cross-Feature Bridge (CFB) further refines multi-level feature fusion. Extensive experiments demonstrate that our VRS-UIE framework sets a new state-of-the-art, delivering superior enhancement performance (surpassing WMamba by 0.89 dB on average) by effectively suppressing water bias and preserving structural and color fidelity. Furthermore, by incorporating efficient convolutional operators and resolution rescaling, we construct a light-weight yet effective scheme, VRS-UIE-S, suitable for real-time UIE applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。