arXiv:2605.07338cs.CV2026-05

构建水下贝类识别数据集,提升生态监测模型的实战能力。

ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs

论文配图:ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs
图 1 · 摘自论文原文
  • 针对真实海底环境设计跨光照、姿态的贝类图像数据集
  • 覆盖32类8691张图,评估80种模型在复杂场景下的表现
  • 含退化模拟测试,适合生态与计算机视觉交叉研究者

全球贝类生物多样性下降威胁沿海生态系统。尽管人工智能在自动化生态监测中展现潜力,但现有海洋底栖数据集常无法适应真实水下环境的复杂性(如光照变化、物种姿态多样),制约视觉模型在实际应用中的泛化能力。为此,我们构建了专为真实生态监测设计的ShellfishNet综合图像基准数据集。该数据集包含32个分类共8,691张图像,通过实地拍摄与网络爬取构建,涵盖复杂真实环境样本,并提供带描述性标题的精选子集。基于此基准,我们系统评估了80种代表性神经网络模型,包括卷积神经网络(CNN)、视觉变压器(ViTs)、状态空间模型(SSMs)及自监督学习(SSL)方法。同时,评估细粒度视觉分类(FGVC)模型性能,并探究主流多模态大语言模型(MLLMs)的图像描述能力。此外,引入图像退化基准测试,模拟浑浊、恶劣天气等常见水下退化场景,评估视觉模型鲁棒性,助力野外生态保育决策的可信性。ShellfishNet致力于为底栖生物智能监测提供数据基础与模型评估标准。

原文摘要 · Abstract (English)

The decline of global shellfish biodiversity poses a severe threat to coastal ecosystems. Although artificial intelligence (AI) technologies show potential for automated ecological monitoring, existing marine benthic datasets often lack adaptation to the complexities of real underwater environments (e.g., variable lighting conditions and diverse species postures), posing challenges for the robust generalization of vision models in practical ecological monitoring. To address this problem, we construct ShellfishNet, a comprehensive image benchmark dataset designed specifically for real-world ecological monitoring constraints. Comprising 8,691 images across 32 taxa, this dataset includes a curated subset annotated with descriptive captions. It is constructed through field photography and web scraping, encompassing samples from complex real-world environments. Based on this benchmark, we systematically evaluate 80 representative neural network models, including Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), State Space Models (SSMs), and Self-Supervised Learning (SSL) methods. Furthermore, we evaluate the performance of fine-grained visual categorization (FGVC) models and investigate the image captioning capabilities of several mainstream multimodal large language models (MLLMs). Meanwhile, we introduce image corruption benchmark tests to simulate common underwater degradation scenarios (turbidity, severe weather) and assess the robustness of vision models, enabling trustworthy decisions on ecological protection in the wild. ShellfishNet is dedicated to providing a data foundation and a model-evaluation benchmark for the intelligent monitoring of benthic organisms.

生态监测图像识别多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。