用模拟退火增强水下密集目标检测的训练数据
A Probabilistic Framework for Improving Dense Object Detection in Underwater Image Data via Annealing-Based Data Augmentation

- 基于模拟退火思想设计新数据增强法,提升图像多样性
- 在深鱼数据集上使YOLOv10检测性能显著提升
- 适合处理真实水下场景中密集、遮挡严重的目标
目标检测模型在受控环境下表现良好,但在真实水下环境中因光照不稳定、能见度低和频繁遮挡导致性能大幅下降。本文针对此问题,提出一种基于伪模拟退火的数据增强框架,以提升复杂水下场景中的检测鲁棒性。利用DeepFish数据集中的分割掩码生成边界框标注,构建定制化检测数据集;通过受Deng等[1]拷贝粘贴策略启发的增强算法,合成逼真的密集鱼类场景。该方法有效提升了训练时的空间多样性和物体密度,从而改善模型对复杂场景的泛化能力。实验表明,该方法显著优于基线YOLOv10模型,尤其在佛罗里达礁石地区实时直播画面中手动标注的挑战性测试集上表现突出。结果证明了该增强策略在真实水下密集场景中提升检测性能的有效性。
原文摘要 · Abstract (English)
Object detection models typically perform well on images captured in controlled environments with stable lighting, water clarity, and viewpoint, but their performance degrades substantially in real-world underwater settings characterized by high variability and frequent occlusions. In this work, we address these challenges by introducing a novel data augmentation framework designed to improve robustness in dense and unconstrained underwater scenes. Using the DeepFish dataset, which contains images of fish in natural environments, we first generate bounding box annotations from provided segmentation masks to construct a custom detection dataset. We then propose a pseudo-simulated annealing-based augmentation algorithm, inspired by the copy-paste strategy of Deng et al. [1], to synthesize realistic crowded fish scenarios. Our approach improves spatial diversity and object density during training, enabling better generalization to complex scenes. Experimental results show that our method significantly outperforms a baseline YOLOv10 model, particularly on a challenging test set of manually annotated images collected from live-stream footage in the Florida Keys. These results demonstrate the effectiveness of our augmentation strategy for improving detection performance in dense, real-world underwater environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。