构建首个统一的大型鱼类检测基准,支持跨水域场景精准识别。
FishDet-M: A Unified Large-Scale Benchmark for Robust Fish Detection and CLIP-Guided Model Selection in Diverse Aquatic Visual Domains
- 整合13个数据集,统一标注格式,实现跨域评估标准化。
- 28种模型实测显示,不同架构在精度与效率间存在显著权衡。
- 基于CLIP的零样本选型框架,动态匹配最佳检测器,适合实时应用。
水下图像中准确的鱼类检测对生态监测、水产养殖自动化和机器人感知至关重要。然而,实际部署受限于数据集碎片化、成像条件多样以及评估协议不一致。为此,我们提出FishDet-M,目前最大的统一鱼类检测基准,涵盖13个公开数据集,覆盖海洋、咸淡水、遮挡及水族馆等多种水生环境。所有数据均采用COCO风格标注,包含边界框和分割掩码,支持一致且可扩展的跨域评估。系统性测试了28种主流目标检测模型,包括YOLOv8至YOLOv12系列、基于R-CNN和DETR的模型。评估使用标准指标mAP、mAP@50、mAP@75,以及尺度相关分析(AP$_S$、AP$_M$、AP$_L$)和延迟、参数量的推理性能分析。结果揭示了不同模型在FishDet-M上表现差异明显,且各架构在精度与效率间存在权衡。为支持自适应部署,引入基于CLIP的模型选择框架,利用视觉-语言对齐动态选出最语义匹配的检测器。该零样本策略无需集成计算即可实现高性能,适用于实时应用。FishDet-M建立了复杂水下场景目标检测的标准化、可复现评估平台。所有数据集、预训练模型和评估工具均开源,助力水下计算机视觉与智能海洋系统研究。
原文摘要 · Abstract (English)
Accurate fish detection in underwater imagery is essential for ecological monitoring, aquaculture automation, and robotic perception. However, practical deployment remains limited by fragmented datasets, heterogeneous imaging conditions, and inconsistent evaluation protocols. To address these gaps, we present \textit{FishDet-M}, the largest unified benchmark for fish detection, comprising 13 publicly available datasets spanning diverse aquatic environments including marine, brackish, occluded, and aquarium scenes. All data are harmonized using COCO-style annotations with both bounding boxes and segmentation masks, enabling consistent and scalable cross-domain evaluation. We systematically benchmark 28 contemporary object detection models, covering the YOLOv8 to YOLOv12 series, R-CNN based detectors, and DETR based models. Evaluations are conducted using standard metrics including mAP, mAP@50, and mAP@75, along with scale-specific analyses (AP$_S$, AP$_M$, AP$_L$) and inference profiling in terms of latency and parameter count. The results highlight the varying detection performance across models trained on FishDet-M, as well as the trade-off between accuracy and efficiency across models of different architectures. To support adaptive deployment, we introduce a CLIP-based model selection framework that leverages vision-language alignment to dynamically identify the most semantically appropriate detector for each input image. This zero-shot selection strategy achieves high performance without requiring ensemble computation, offering a scalable solution for real-time applications. FishDet-M establishes a standardized and reproducible platform for evaluating object detection in complex aquatic scenes. All datasets, pretrained models, and evaluation tools are publicly available to facilitate future research in underwater computer vision and intelligent marine systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。