统一修复与检测,让水下显著目标更准更真。
UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection

- 修复与检测联合优化,用语义指导图像恢复
- 在多个数据集上超越现有方法,提升显著目标识别精度
- 适合水下视觉、海洋探测等场景的科研与应用
水下显著目标检测(USOD)在海洋视觉任务中至关重要,但因选择性吸收和介质散射导致严重视觉退化而极具挑战。传统方法采用“先增强后检测”的串行流程,常因低层修复与高层语义分离,造成语义不一致,修复结果未必利于检测,甚至引入无关噪声。为此,我们提出UniV2D——一种统一视觉到检测的网络,通过相互促进的框架联合优化视觉修复与显著目标检测。不同于依赖独立流程或刚性物理先验的方法,UniV2D采用语义驱动学习:高层显著性语义主动引导修复过程,修复后的视觉线索又反向增强语义感知。具体地,该模型采用分层双分支架构:首先通过自校准解码器预测初始显著性掩码,并结合掩码感知修复模块重建图像内容;随后利用带有跨层级调制的显著性引导精修模块,对齐结构保真度与语义一致性。大量实验表明,UniV2D在多个基准测试中显著优于现有最先进方法,在定量与定性评估上均建立新标准,推动水下联合感知发展。
原文摘要 · Abstract (English)
Underwater salient object detection (USOD) plays a vital role in marine vision tasks but remains fundamentally challenging due to severe visual degradation, such as selective absorption and medium scattering. Conventional pipelines typically adopt a sequential "enhance-then-detect" paradigm. However, isolating low-level visual restoration from high-level semantic perception often leads to semantic inconsistency, where the restored images may not be optimal for detection and can even introduce task-irrelevant noise. To break this sequential bottleneck, we propose UniV2D, a Unified Vision-to-Detection Network that jointly optimizes visual restoration and salient object detection within a mutually beneficial framework. Unlike traditional methods that rely on disjointed pipelines or rigid physical priors, UniV2D introduces a semantic-driven learning paradigm: high-level saliency semantics actively guide the restoration process, while the restored visual cues reciprocally enhance saliency perception. Specifically, UniV2D features a hierarchical dual-branch architecture. It first employs a self-calibrated decoder to predict initial saliency masks alongside a mask-aware restoration module to reconstruct image content. Subsequently, a saliency-guided refinement module equipped with cross-level modulation is utilized to align structural fidelity with semantic consistency. Extensive experiments across multiple benchmarks demonstrate that UniV2D significantly outperforms state-of-the-art methods in both quantitative and qualitative evaluations, establishing a new standard for joint underwater perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。