arXiv:2503.08152cs.CV2025-03被引 3

用深度信息和运动差异编码,提升水下模糊物体的计数精度。

Depth-Assisted Network for Indiscernible Marine Object Counting with Adaptive Motion-Differentiated Feature Encoding

  • 引入深度辅助与动态运动加权,增强复杂水下场景特征表达。
  • 在自建50视频、40800标注点数据集上达到领先性能。
  • 适合水下监测、海洋生物计数等需要高精度弱可见目标识别的研究者。

水下模糊物体计数面临能见度低、物体重叠遮挡以及背景与前景外观、颜色、纹理高度相似等挑战。为解决视频类数据稀缺问题,我们构建了一个包含50段视频的新数据集,从中提取约800帧并标注了约40,800个点状目标,真实还原水下环境中对象与背景难以区分的复杂情况。为此,提出一种深度辅助网络,包含主干编码模块及三个分支:深度辅助分支、密度估计分支、运动权重生成分支。深度辅助分支提取的深度感知特征经深度增强编码器强化,提升物体表征能力;运动权重生成分支输出的权重用于优化自适应流估计模块中的多尺度感知特征。实验表明,该方法不仅在所提数据集上达到当前最优表现,还在三个额外的视频人群计数数据集上取得具有竞争力的结果。预训练模型、代码与数据集已公开于https://github.com/OUCVisionGroup/VIMOC-Net。

原文摘要 · Abstract (English)

Indiscernible marine object counting encounters numerous challenges, including limited visibility in underwater scenes, mutual occlusion and overlap among objects, and the dynamic similarity in appearance, color, and texture between the background and foreground. These factors significantly complicate the counting process. To address the scarcity of video-based indiscernible object counting datasets, we have developed a novel dataset comprising 50 videos, from which approximately 800 frames have been extracted and annotated with around 40,800 point-wise object labels. This dataset accurately represents real underwater environments where indiscernible marine objects are intricately integrated with their surroundings, thereby comprehensively illustrating the aforementioned challenges in object counting. To address these challenges, we propose a depth-assisted network with adaptive motion-differentiated feature encoding. The network consists of a backbone encoding module and three branches: a depth-assisting branch, a density estimation branch, and a motion weight generation branch. Depth-aware features extracted by the depth-assisting branch are enhanced via a depth-enhanced encoder to improve object representation. Meanwhile, weights from the motion weight generation branch refine multi-scale perception features in the adaptive flow estimation module. Experimental results demonstrate that our method not only achieves state-of-the-art performance on the proposed dataset but also yields competitive results on three additional video-based crowd counting datasets. The pre-trained model, code, and dataset are publicly available at https://github.com/OUCVisionGroup/VIMOC-Net.

水下计数深度感知运动编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。