用光学去噪方法提升声呐图像检测效果,提出多源融合新框架。
Can Optical Denoising Clean Sonar Images? A Benchmark and Fusion Approach
- 对比九种深度去噪模型在声呐图像上的表现。
- 去噪后目标检测平均提升12.3%准确率,但效果因模型而异。
- 提出像素级互监督融合框架,实现协同去噪增效。
声呐图像中的目标检测对水下机器人自主导航与资源勘探至关重要。然而,声呐图像固有的复杂噪声(如斑点、混响及非高斯噪声)严重降低检测精度。尽管光学图像去噪技术已取得显著进展,其在声呐数据上的适用性仍缺乏系统研究。本研究首次系统评估了九种前沿深度去噪模型(包括Neighbor2Neighbor、Blind2Unblind和DSPNet)在声呐图像预处理中的表现。基于五个公开声呐数据集,评估其对四种代表性检测算法(YOLOX、Faster R-CNN、SSD300、SSDMobileNetV2)的影响。实验回答三个关键问题:一、光学去噪架构能否有效迁移至声呐数据;二、哪些模型家族在声呐噪声下表现更优;三、去噪是否真正提升实际检测性能。结果表明,去噪普遍提升检测性能,但不同方法因对特定噪声类型存在固有偏好而表现不一。为融合互补优势,我们提出一种像素级互监督的多源去噪融合框架,各去噪器输出相互监督,形成协同优化机制,生成更清晰图像。
原文摘要 · Abstract (English)
Object detection in sonar images is crucial for underwater robotics applications including autonomous navigation and resource exploration. However, complex noise patterns inherent in sonar imagery, particularly speckle, reverberation, and non-Gaussian noise, significantly degrade detection accuracy. While denoising techniques have achieved remarkable success in optical imaging, their applicability to sonar data remains underexplored. This study presents the first systematic evaluation of nine state-of-the-art deep denoising models with distinct architectures, including Neighbor2Neighbor with varying noise parameters, Blind2Unblind with different noise configurations, and DSPNet, for sonar image preprocessing. We establish a rigorous benchmark using five publicly available sonar datasets and assess their impact on four representative detection algorithms: YOLOX, Faster R-CNN, SSD300, and SSDMobileNetV2. Our evaluation addresses three unresolved questions: first, how effectively optical denoising architectures transfer to sonar data; second, which model families perform best against sonar noise; and third, whether denoising truly improves detection accuracy in practical pipelines. Extensive experiments demonstrate that while denoising generally improves detection performance, effectiveness varies across methods due to their inherent biases toward specific noise types. To leverage complementary denoising effects, we propose a mutually-supervised multi-source denoising fusion framework where outputs from different denoisers mutually supervise each other at the pixel level, creating a synergistic framework that produces cleaner images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。