arXiv:2510.14594cs.CV2025-10被引 1

用视觉模型融合技术,把动物分类结果从属级提升到物种级。

Hierarchical Re-Classification: Combining Animal Classification Models with Vision Transformers

  • 分五步融合EfficientNet与CLIP,优化高阶分类标签
  • 在4018张图像上实现96.5%的重分类准确率
  • 适合野生动物监测与生态研究场景

当前先进的动物分类模型如SpeciesNet可识别数千个物种,但采用保守的汇总策略,导致大量动物仅被标注至高阶分类层级。本文提出一种用于Animal Detect平台的分层再分类系统,结合SpeciesNet EfficientNetV2-M预测结果、CLIP嵌入与度量学习,将高阶分类标签细化至物种级别。该系统包含五个阶段:高置信度接受、鸟类优先覆盖、中心点构建、三元组损失度量学习及自适应余弦距离评分。在LILA BC沙漠狮保护数据集的一个子集(4,018张图像,15,031次检测)上评估,成功从“空白”和“动物”标签中恢复761次鸟类检测,并对456次原标为“动物”“哺乳类”或“空白”的检测进行重分类,准确率达96.5%,实现64.9%的物种级识别率。

原文摘要 · Abstract (English)

State-of-the-art animal classification models like SpeciesNet provide predictions across thousands of species but use conservative rollup strategies, resulting in many animals labeled at high taxonomic levels rather than species. We present a hierarchical re-classification system for the Animal Detect platform that combines SpeciesNet EfficientNetV2-M predictions with CLIP embeddings and metric learning to refine high-level taxonomic labels toward species-level identification. Our five-stage pipeline (high-confidence acceptance, bird override, centroid building, triplet-loss metric learning, and adaptive cosine-distance scoring) is evaluated on a segment of the LILA BC Desert Lion Conservation dataset (4,018 images, 15,031 detections). After recovering 761 bird detections from "blank" and "animal" labels, we re-classify 456 detections labeled animal, mammal, or blank with 96.5% accuracy, achieving species-level identification for 64.9 percent

动物识别视觉模型分类优化生态监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。