用多目标追踪提升相机陷阱中的动物识别一致性
Multi-Object Tracking Consistently Improves Wildlife Inference

- 用标准多目标追踪模型关联连续帧检测结果
- 融合分类置信度,使单个个体标签更稳定
- 在三个数据集上显著提升分类准确率,适合生态监测应用
相机陷阱已成为生态研究和生物多样性保护中常见的野生动物监测工具。尽管野生动物分类模型在高质量标注数据集上表现优异,但在真实环境下的性能仍受限制,尤其在处理时间连贯的数据序列时会出现预测不一致的问题——同一个体的标签在不同帧间频繁跳变。本研究利用相机陷阱数据的时间特性,通过引入多种标准多目标追踪(MOT)模型,将连续帧中的检测结果进行关联,生成清晰的个体轨迹。基于这些轨迹,对分类模型输出的softmax概率进行融合,生成统一的共识类别预测,从而覆盖由噪声引起的误判。实验结果表明,该策略在所有数据集和各项指标上均优于独立分类器。其中,表现最佳的MOT模型在三个数据集上分别实现5.1%、3.1%和2.0%的加权F1分数提升。
原文摘要 · Abstract (English)
Camera traps have become a common tool for wildlife monitoring efforts in ecological research and biodiversity conservation. Wildlife classification models have benefited from the increase in wildlife visual data. These models reach high levels of accuracy on curated, high-quality datasets. However, their performance remains sensitive to real-world environmental constraints. They often produce inconsistent predictions when performing inference on temporally coherent sequences. The predicted label for a single individual shifts rapidly between frames. This study exploits the temporal nature of camera-trap data to augment inferred predictions from a wildlife classification model. Specifically, we adopt several standard Multi-Object Tracking (MOT) models to link detections across consecutive frames. The curated trajectories are used to fuse the softmax class probabilities. The fused probability score produces a single consensus class label estimate that overrides misclassifications caused by noise. The analysis of the experimental results shows that our proposed strategy improves over a standalone classifier over all datasets and for each metric. Specifically, the best-performing MOT models gain a weighted F1-Score of 5.1%, 3.1% and 2.0% over the classifier across three MOT datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。