arXiv:2510.25797cs.CVcs.CL2025-10

通过时空建模与注意力机制,提升水下目标检测精度。

Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks

  • 引入时序增强模块和空间注意力,改进YOLOv5模型
  • 新模型mAP@50-95达0.813,显著优于原版的0.563
  • 适合复杂动态水下环境,尤其抗遮挡与运动模糊

本研究评估了时空建模与空间注意力机制在深度学习模型中对水下目标检测的效果。首先对比标准YOLOv5与时序增强版本T-YOLOv5的性能;随后在T-YOLOv5基础上引入卷积块注意力模块(CBAM),构建T-YOLOv5 with CBAM。实验结果表明,标准YOLOv5的mAP@50-95为0.563,而T-YOLOv5提升至0.813,加入CBAM后为0.811,均显著优于原始模型。该方法在突发运动、部分遮挡和渐进运动等复杂场景中表现出更强的检测准确性和泛化能力。尽管在简单场景中精度略有下降,但整体提升了水下动态环境下的检测可靠性。

原文摘要 · Abstract (English)

This study examines the effectiveness of spatio-temporal modeling and the integration of spatial attention mechanisms in deep learning models for underwater object detection. Specifically, in the first phase, the performance of temporal-enhanced YOLOv5 variant T-YOLOv5 is evaluated, in comparison with the standard YOLOv5. For the second phase, an augmented version of T-YOLOv5 is developed, through the addition of a Convolutional Block Attention Module (CBAM). By examining the effectiveness of the already pre-existing YOLOv5 and T-YOLOv5 models and of the newly developed T-YOLOv5 with CBAM. With CBAM, the research highlights how temporal modeling improves detection accuracy in dynamic marine environments, particularly under conditions of sudden movements, partial occlusions, and gradual motion. The testing results showed that YOLOv5 achieved a mAP@50-95 of 0.563, while T-YOLOv5 and T-YOLOv5 with CBAM outperformed with mAP@50-95 scores of 0.813 and 0.811, respectively, highlighting their superior accuracy and generalization in detecting complex objects. The findings demonstrate that T-YOLOv5 significantly enhances detection reliability compared to the standard model, while T-YOLOv5 with CBAM further improves performance in challenging scenarios, although there is a loss of accuracy when it comes to simpler scenarios.

水下检测时空建模注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。