首个针对相机陷阱物种识别随时间演化的统一研究,揭示长期部署中的关键挑战。
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
- 构建546个相机陷阱的流式评估基准,模拟真实时间序列部署场景。
- 模型更新常导致性能下降,因物种分布与背景变化剧烈且类别极不平衡。
- 融合模型更新与后处理可显著提效,但距离最优仍有差距,适合生态监测实践者参考。
相机陷阱对大规模生物多样性监测至关重要,但自动化分析仍面临挑战,主要源于部署环境多样。尽管计算机视觉领域多将其视为跨域泛化问题,生态实践者更关注固定站点随时间的可靠识别——生态系统动态导致背景和动物分布产生深刻的时间漂移。为此,我们首次开展相机陷阱物种识别随时间演化的统一研究。引入一个包含546个相机陷阱的真实基准,采用时间顺序划分的流式评估协议。端用户导向的研究得出四项关键发现:(1)生物基础模型(如BioCLIP 2)在多数站点初始阶段即表现不佳,凸显站点适配必要性;(2)在真实评估下(用历史数据更新模型并评估未来区间),简单适配反而可能使性能低于零样本表现;(3)困难根源在于严重类别不平衡及连续区间间物种分布与背景的显著时间漂移;(4)有效结合模型更新与后处理技术可大幅提升准确率,但仍存在与上限的差距。最后,我们提出若干关键开放问题,如预测零样本模型在新站点的成功概率,以及何时需要更新模型。本基准与分析为生态实践提供可操作部署指南,同时为视觉与机器学习研究开辟新方向。
原文摘要 · Abstract (English)
Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment environments. While the computer vision community has mostly framed this challenge as cross-domain generalization, this perspective overlooks a primary challenge faced by ecological practitioners: maintaining reliable recognition at the fixed site over time, where the dynamic nature of ecosystems introduces profound temporal shifts in both background and animal distributions. To bridge this gap, we present the first unified study of camera-trap species recognition over time. We introduce a realistic benchmark comprising 546 camera traps with a streaming protocol that evaluates models over chronologically ordered intervals. Our end-user-centric study yields four key findings. (1) Biological foundation models (e.g., BioCLIP 2) underperform at numerous sites even in initial intervals, underscoring the necessity of site-specific adaptation. (2) Adaptation is challenging under realistic evaluation: when models are updated using past data and evaluated on future intervals (mirrors real deployment lifecycles), naive adaptation can even degrade below zero-shot performance. (3) We identify two drivers of this difficulty: severe class imbalance and pronounced temporal shift in both species distribution and backgrounds between consecutive intervals. (4) We find that effective integration of model-update and post-processing techniques can largely improve accuracy, though a gap from the upper bounds remains. Finally, we highlight critical open questions, such as predicting when zero-shot models will succeed at a new site and determining whether/when model updates are necessary. Our benchmark and analysis provide actionable deployment guidelines for ecological practitioners while establishing new directions for future research in vision and machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。