arXiv:2512.05996cs.CVcs.CY2025-12

用弱监督实现水下鱼群检测分割计数,性能显著提升。

FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting

  • 设计检测-计数提示,保证空间一致性
  • 弱监督下AP提升20%,计数误差降35%
  • 适合海洋生态监测与标注成本高的场景

分析水下鱼类图像对生态监测至关重要,但受视觉退化和标注成本高等因素制约。本文提出FishDetector-R1,一种基于多模态大模型的统一框架,可在弱监督条件下完成鱼类检测、分割与计数。在DeepFish数据集上,该框架相比基线模型显著提升:平均精度(AP)提高20%,交并比(mIoU)提升10%,平均绝对误差(MAE)降低30%,游戏评估指标(GAME)下降35%。核心贡献包括:一种新型检测-计数提示机制,确保检测与计数的空间一致性;以及基于可验证奖励的强化学习(RLVR),结合稀疏点标注实现可扩展训练。消融实验验证了奖励设计的有效性。性能提升在其他水下数据集上也具有良好的泛化能力,表明其强跨域鲁棒性。FishDetector-R1为弱监督下的海洋视觉理解提供了可靠且可扩展的解决方案。

原文摘要 · Abstract (English)

Analyzing underwater fish imagery is critical for ecological monitoring but remains difficult due to visual degradation and costly annotations. We introduce FishDetector-R1, a unified MLLM-based framework for fish detection, segmentation, and counting under weak supervision. On the DeepFish dataset, our framework achieves substantial gains over baselines, improving AP by 20% and mIoU by 10%, while reducing MAE by 30% and GAME by 35%. These improvements stem from two key components: a novel detect-to-count prompt that enforces spatially consistent detections and counts, and Reinforcement Learning from Verifiable Reward (RLVR) with a complementary scalable paradigm leveraging sparse point labels. Ablation studies further validate the effectiveness of this reward design. Moreover, the improvement generalizes well to other underwater datasets, confirming strong cross-domain robustness. Overall, FishDetector-R1 provides a reliable and scalable solution for accurate marine visual understanding via weak supervision. The project page for FishDetector-R1 is https://umfieldrobotics.github.io/FishDetector-R1.

弱监督目标检测水下视觉多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。