arXiv:2503.11341cs.CV2025-03CVPR被引 7

用自监督预训练提升微粒浮游生物识别精度,减少标注依赖。

Self-Supervised Pretraining for Fine-Grained Plankton Recognition

  • 通过掩码自编码在海量浮游生物图像上进行自监督预训练
  • 小样本下识别准确率显著高于传统ImageNet预训练方法
  • 适用于新数据集快速适配,尤其适合标注资源有限的研究场景

浮游生物识别是计算机视觉中的重要问题,因其在海洋食物链和碳捕获中的关键作用,亟需物种级监测。然而,该任务因细粒度特征及成像设备、物种分布差异导致的数据分布偏移而具有挑战性。随着浮游生物图像数据集以越来越快的速度积累,亟需无需大量人工标注的通用识别模型。本文研究大规模自监督预训练在细粒度浮游生物识别中的应用。首先,利用掩码自编码与大量多样化浮游生物图像数据预训练一个通用浮游生物图像编码器;随后通过微调,在仅需极少标注图像的情况下,为新数据集构建高精度识别模型。实验表明,在训练数据有限时,使用多样化浮游生物数据进行自监督预训练可显著提升识别准确率,优于标准ImageNet预训练;若在预训练阶段引入未标注的目标数据,性能还可进一步提升。

原文摘要 · Abstract (English)

Plankton recognition is an important computer vision problem due to plankton's essential role in ocean food webs and carbon capture, highlighting the need for species-level monitoring. However, this task is challenging due to its fine-grained nature and dataset shifts caused by different imaging instruments and varying species distributions. As new plankton image datasets are collected at an increasing pace, there is a need for general plankton recognition models that require minimal expert effort for data labeling. In this work, we study large-scale self-supervised pretraining for fine-grained plankton recognition. We first employ masked autoencoding and a large volume of diverse plankton image data to pretrain a general-purpose plankton image encoder. Then we utilize fine-tuning to obtain accurate plankton recognition models for new datasets with a very limited number of labeled training images. Our experiments show that self-supervised pretraining with diverse plankton data clearly increases plankton recognition accuracy compared to standard ImageNet pretraining when the amount of training data is limited. Moreover, the accuracy can be further improved when unlabeled target data is available and utilized during the pretraining.

浮游生物识别自监督学习小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。