arXiv:2409.19323cs.RO2024-09被引 1

用轻量Transformer实现高密度相似鱼群的快速精准检测

Intelligent Fish Detection System with Similarity-Aware Transformer

  • 设计相似性感知多级编码器,提升不同大小鱼的特征区分度
  • 引入软阈值注意力机制,有效去除背景噪声并保留边缘细节
  • 在85个真实场景视频上实现超80帧/秒,适合工业部署

水陆转移场景中的鱼类检测对渔业具有重要意义。然而,人工协作检测效率低、成本高且精度不足。为此,本文设计了一种轻量级、即插即用的边缘智能视觉系统,结合高速相机实现自动快速鱼类检测。提出一种新型相似性感知视觉Transformer(FishViT),用于识别密集且外观相似的鱼群。具体地,构建了相似性感知多级编码器,平行增强多尺度特征,获得不同尺寸鱼的判别性表征;同时引入软阈值注意力机制,有效消除图像背景噪声,并准确捕捉不同相似鱼类的边缘细节与整体特征。收集了85个高帧率、高分辨率的真实水陆转移视频序列,建立基准测试集。基于该挑战性基准的全面评估验证了FishViT的鲁棒性与有效性,实测速度超过80 FPS。实际工作场景测试进一步证实了方法的实用性。代码与演示视频已开源。

原文摘要 · Abstract (English)

Fish detection in water-land transfer has significantly contributed to the fishery. However, manual fish detection in crowd-collaboration performs inefficiently and expensively, involving insufficient accuracy. To further enhance the water-land transfer efficiency, improve detection accuracy, and reduce labor costs, this work designs a new type of lightweight and plug-and-play edge intelligent vision system to automatically conduct fast fish detection with high-speed camera. Moreover, a novel similarity-aware vision Transformer for fast fish detection (FishViT) is proposed to onboard identify every single fish in a dense and similar group. Specifically, a novel similarity-aware multi-level encoder is developed to enhance multi-scale features in parallel, thereby yielding discriminative representations for varying-size fish. Additionally, a new soft-threshold attention mechanism is introduced, which not only effectively eliminates background noise from images but also accurately recognizes both the edge details and overall features of different similar fish. 85 challenging video sequences with high framerate and high-resolution are collected to establish a benchmark from real fish water-land transfer scenarios. Exhaustive evaluation conducted with this challenging benchmark has proved the robustness and effectiveness of FishViT with over 80 FPS. Real work scenario tests validate the practicality of the proposed method. The code and demo video are available at https://github.com/vision4robotics/FishViT.

目标检测视觉Transformer工业视觉轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。