arXiv:2608.18246cs.CVcs.AI2026-08

用视觉提示统一检测与识别,提升野生动物细粒度辨识效果

Visual-Prompt Guided Wildlife Instance-Level Recognition

论文配图:Visual-Prompt Guided Wildlife Instance-Level Recognition
图 1 · 摘自论文原文
  • 单阶段端到端模型,在潜在空间中直接搜索身份
  • 在Wildlife-100数据集上达到30.584%的mAP,接近两阶段方法的70%
  • 适合需要快速部署的野生动物监测场景

细粒度野生动物重识别仍是研究难点。现有最优方法采用检测与重识别分步流程。本文提出一种单阶段端到端检测与重识别模型,可在潜在空间中进行身份搜索。采用DINOv2获取鲁棒的空间几何特征,MegaDescriptor用于野生动物重识别。通过提示增强潜在查询特征,检测解码器在场景潜在空间中定位目标身份边界。初步结果表明,该模型在Wildlife-100数据集上取得30.584%的平均精度均值,相较于当前最优两阶段方法的44.89%仍具竞争力。定性结果展示出对动物身份的有效框选与识别。

原文摘要 · Abstract (English)

Fine-grained wildlife re-identification remains a challenging area in research. Current state-of-the-art approaches apply a detection and re-identification pipeline. We propose a one-stage end-to-end detection and re-identification model that performs identity searching within the latent space. We adopt DINOv2 for robust spatial geometry and MegaDescriptor for wildlife re-identification. We enhance latent queries with prompt re-identification features. A detection decoder queries the scene latent space to establish object boundaries around the target identity. Preliminary findings reflect a competitive mean average precision score of 30.584% compared to the state-of-the-art two stage approach of 44.89%. Qualitative results depict effective bounding and identification of animal identities.

野生动物识别细粒度识别单阶段模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。