AQUA20数据集助力水下物种识别,应对浑浊光照等挑战。
AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions
- 构建8171张图像的水下物种数据集,覆盖20类真实复杂环境样本。
- ConvNeXt模型在该数据集上达到90.69%准确率与88.92% F1分数。
- 提供可视化分析工具,帮助理解模型优劣,适合水下视觉研究者。
由于浑浊、低照度和遮挡等复杂失真,水下环境中的鲁棒视觉识别仍是重大挑战,严重削弱了标准视觉系统的性能。本文提出AQUA20,一个包含8,171张水下图像的综合性基准数据集,涵盖20种海洋生物,反映真实环境下的光照、浑浊、遮挡等挑战,为水下视觉理解提供宝贵资源。评估了13种先进深度学习模型,包括轻量级CNN(SqueezeNet、MobileNetV2)和基于Transformer的架构(ViT、ConvNeXt),结果表明ConvNeXt表现最佳:Top-3准确率达98.82%,Top-1准确率为90.69%,整体F1分数达88.92%,参数量适中。其他模型也揭示了复杂性与性能之间的权衡。通过GRAD-CAM和LIME进行可解释性分析,揭示了模型的优势与缺陷。结果表明水下物种识别仍有巨大提升空间,证实AQUA20是未来研究的重要基础。数据集已公开:https://huggingface.co/datasets/taufiktrf/AQUA20。
原文摘要 · Abstract (English)
Robust visual recognition in underwater environments remains a significant challenge due to complex distortions such as turbidity, low illumination, and occlusion, which severely degrade the performance of standard vision systems. This paper introduces AQUA20, a comprehensive benchmark dataset comprising 8,171 underwater images across 20 marine species reflecting real-world environmental challenges such as illumination, turbidity, occlusions, etc., providing a valuable resource for underwater visual understanding. Thirteen state-of-the-art deep learning models, including lightweight CNNs (SqueezeNet, MobileNetV2) and transformer-based architectures (ViT, ConvNeXt), were evaluated to benchmark their performance in classifying marine species under challenging conditions. Our experimental results show ConvNeXt achieving the best performance, with a Top-3 accuracy of 98.82% and a Top-1 accuracy of 90.69%, as well as the highest overall F1-score of 88.92% with moderately large parameter size. The results obtained from our other benchmark models also demonstrate trade-offs between complexity and performance. We also provide an extensive explainability analysis using GRAD-CAM and LIME for interpreting the strengths and pitfalls of the models. Our results reveal substantial room for improvement in underwater species recognition and demonstrate the value of AQUA20 as a foundation for future research in this domain. The dataset is publicly available at: https://huggingface.co/datasets/taufiktrf/AQUA20.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。