arXiv:2410.12769cs.CV2024-10ECCV被引 10

用零样本模型解决相机陷阱图像分类中的地域过拟合问题

Towards Zero-Shot Camera Trap Image Categorization

  • 结合MegaDetector与双分类器,减少地域特异性偏差
  • 在多个数据集上相对误差降低42%至75%,新地点准确率提升一半
  • 基于DINOv2和FAISS的零样本方案表现接近有监督方法,适合无标注场景

本文探索相机陷阱图像自动分类的新方法。首先,使用单一模型对所有图像进行基准测试;其次,评估MegaDetector与一个或多个分类器、Segment Anything结合的效果,以缓解位置特异性过拟合;最后,提出并测试两种基于大语言模型和基础模型(如DINOv2、BioCLIP、BLIP、ChatGPT)的零样本方案。在两个公开数据集(新西兰的WCT、美国西南部的CCT20)和一个私有数据集(中欧的CEF)上的评估表明,将MegaDetector与两个独立分类器结合的方法达到最高准确率。该方法在CCT20上使BEiTV2单模型相对误差降低约42%,在CEF上降低48%,在WCT上降低75%。此外,去除背景后,新地点的错误率减半。基于DINOv2和FAISS的零样本流程在CCT20和CEF上分别取得1.0%和4.7%的精度优势,凸显了零样本方法在相机陷阱图像分类中的潜力。

原文摘要 · Abstract (English)

This paper describes the search for an alternative approach to the automatic categorization of camera trap images. First, we benchmark state-of-the-art classifiers using a single model for all images. Next, we evaluate methods combining MegaDetector with one or more classifiers and Segment Anything to assess their impact on reducing location-specific overfitting. Last, we propose and test two approaches using large language and foundational models, such as DINOv2, BioCLIP, BLIP, and ChatGPT, in a zero-shot scenario. Evaluation carried out on two publicly available datasets (WCT from New Zealand, CCT20 from the Southwestern US) and a private dataset (CEF from Central Europe) revealed that combining MegaDetector with two separate classifiers achieves the highest accuracy. This approach reduced the relative error of a single BEiTV2 classifier by approximately 42\% on CCT20, 48\% on CEF, and 75\% on WCT. Besides, as the background is removed, the error in terms of accuracy in new locations is reduced to half. The proposed zero-shot pipeline based on DINOv2 and FAISS achieved competitive results (1.0\% and 4.7\% smaller on CCT20, and CEF, respectively), which highlights the potential of zero-shot approaches for camera trap image categorization.

零样本学习图像分类动物监测跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。