arXiv:2604.27538cs.CV2026-04

针对植物识别优化自监督学习,提升细粒度分类效果。

Self-Supervised Learning of Plant Image Representations

论文配图:Self-Supervised Learning of Plant Image Representations
图 1 · 摘自论文原文
  • 用仿射变换和海报化替代传统增强,更适配植物图像特征。
  • 在iNaturalist植物子集上训练的SimDINOv2表现优于ImageNet-1K模型。
  • 少样本场景下可超越监督基线,适合生物多样性监测应用。

自动化植物识别对生物多样性监测至关重要,但现有方法依赖专家标注数据,受限于监督学习。自监督学习(SSL)提供可扩展替代方案,但现有方法多为粗粒度视觉任务设计,难以迁移至植物物种识别等细粒度领域。本文研究植物图像表示学习中的自监督方法,发现常用增强如高斯模糊、灰度化和太阳化会移除植物图像中细微的判别线索,不利于细粒度识别。我们提出采用仿射变换和海报化作为更合适的增强方式。进一步表明,在iNaturalist 2021 Plantae子集上训练SimDINOv2,其表示能力显著优于在ImageNet-1K上训练的模型,凸显领域特定数据的重要性。该结论在ViT-Base与ViT-Large架构上均成立。此外,模型在少样本下游任务中表现优异,部分情况下超越强监督基线Pl@ntCLEF与BioCLIP。结果强调了领域适配的增强策略与数据选择在自监督学习中的关键作用,为构建可扩展的生物多样性监测模型提供了实用指导。

原文摘要 · Abstract (English)

Automated plant recognition plays a crucial role in biodiversity monitoring and conservation, yet current approaches rely heavily on supervised learning, which is limited by the availability of expert-labeled data. Self-supervised learning (SSL) offers a scalable alternative, but existing methods and training protocols are largely designed for coarse-grained visual tasks and may not transfer well to fine-grained domains such as plant species recognition. In this work, we investigate SSL for plant image representation learning. We show that commonly used augmentations in SSL pipelines - such as Gaussian blur, grayscale conversion, and solarization - are detrimental in the context of plant images, as they remove subtle discriminative cues essential for fine-grained recognition. We instead identify alternative transformations, including affine and posterization, that are better suited to this domain. We further demonstrate that training SimDINOv2 on the iNaturalist 2021 Plantae subset yields significantly stronger representations than training on ImageNet-1K, highlighting the importance of domain-specific data for SSL. Our findings are consistent across both ViT-Base and ViT-Large architectures. Moreover, our models achieve competitive performance and sometimes outperform strong supervised baselines Pl@ntCLEF and BioCLIP on downstream plant recognition tasks in few-shot settings. Overall, our results highlight the critical importance of domain-adapted augmentation strategies and dataset selection in self-supervised learning, and provide practical guidelines for building scalable models for biodiversity monitoring.

自监督学习植物识别细粒度分类图像增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。