arXiv:2510.14596cs.CV2025-10被引 1

用视觉变压器零样本整理野生动物照片,提升物种识别效率。

Zero-Shot Wildlife Sorting Using Vision Transformers: Evaluating Clustering and Continuous Similarity Ordering

  • 用自监督视觉变压器提取特征,结合聚类与降维技术组织无标签图像。
  • 在5个物种上达88.6%准确率,鱼类排序一致性高达95.2%。
  • 适合需要快速分析海量相机陷阱图像的研究者使用。

相机陷阱生成数百万张野生动物图像,但许多数据集包含现有分类器未覆盖的物种。本文在Animal Detect平台中评估了基于自监督视觉变压器的零样本方法,用于组织无标签野生动物图像。比较了三种架构(CLIP、DINOv2、MegaDescriptor)与无监督聚类方法(DBSCAN、GMM)结合主成分分析(PCA)、UMAP的性能,并通过t-SNE实现连续1维相似性排序。在包含5个物种的测试集上(仅用于评估),使用DINOv2+UMAP+GMM的方法达到88.6%准确率(宏平均F1=0.874);1D排序对哺乳动物和鸟类的连贯性达88.2%,鱼类为95.2%(共1,500张图像)。基于结果,已将连续相似性排序部署至生产环境,支持快速探索分析并加速人工标注流程,助力生物多样性监测。

原文摘要 · Abstract (English)

Camera traps generate millions of wildlife images, yet many datasets contain species that are absent from existing classifiers. This work evaluates zero-shot approaches for organizing unlabeled wildlife imagery using self-supervised vision transformers, developed and tested within the Animal Detect platform for camera trap analysis. We compare unsupervised clustering methods (DBSCAN, GMM) across three architectures (CLIP, DINOv2, MegaDescriptor) combined with dimensionality reduction techniques (PCA, UMAP), and we demonstrate continuous 1D similarity ordering via t-SNE projection. On a 5-species test set with ground truth labels used only for evaluation, DINOv2 with UMAP and GMM achieves 88.6 percent accuracy (macro-F1 = 0.874), while 1D sorting reaches 88.2 percent coherence for mammals and birds and 95.2 percent for fish across 1,500 images. Based on these findings, we deployed continuous similarity ordering in production, enabling rapid exploratory analysis and accelerating manual annotation workflows for biodiversity monitoring.

零样本视觉变压器生物多样性聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。