首个量化背景对野生动物行为识别影响的数据集,助力模型泛化能力研究。
The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recognition
- 构建双视频配对数据集,分离动物与背景信息以分析其影响
- 提出隐空间归一化方法,使模型在分布外场景下准确率提升超5%
- 揭示背景时长等新因素对模型性能的关键作用,适合视觉算法与保护生物学者
基于相机陷阱视频的计算机视觉分析对野生动物保护至关重要,因行为变化是种群健康早期预警信号。尽管已有多个高质量动物行为数据集和方法发布,但行为相关背景信息的作用及其对分布外泛化的影响仍不明晰。为此,我们推出PanAf-FGBG数据集,包含超过20小时的野生黑猩猩行为视频,覆盖350余个独立摄像头位置。每个带黑猩猩的前景视频均配有同一位置无黑猩猩的背景视频,形成双重视图:重叠与互斥摄像头设置。这首次实现对分布内与分布外条件的直接评估,并可量化背景对行为识别模型的影响。所有片段附有丰富行为标注与元数据(包括唯一摄像头ID与详细场景描述)。我们建立了多个基准,并提出一种高效的隐空间归一化技术,在卷积与基于变换器的模型上分别将分布外mAP提升5.42%和3.75%。最后,深入分析背景在分布外识别中的作用,首次揭示背景持续时间(即前景视频中背景帧数)的重要影响。
原文摘要 · Abstract (English)
Computer vision analysis of camera trap video footage is essential for wildlife conservation, as captured behaviours offer some of the earliest indicators of changes in population health. Recently, several high-impact animal behaviour datasets and methods have been introduced to encourage their use; however, the role of behaviour-correlated background information and its significant effect on out-of-distribution generalisation remain unexplored. In response, we present the PanAf-FGBG dataset, featuring 20 hours of wild chimpanzee behaviours, recorded at over 350 individual camera locations. Uniquely, it pairs every video with a chimpanzee (referred to as a foreground video) with a corresponding background video (with no chimpanzee) from the same camera location. We present two views of the dataset: one with overlapping camera locations and one with disjoint locations. This setup enables, for the first time, direct evaluation of in-distribution and out-of-distribution conditions, and for the impact of backgrounds on behaviour recognition models to be quantified. All clips come with rich behavioural annotations and metadata including unique camera IDs and detailed textual scene descriptions. Additionally, we establish several baselines and present a highly effective latent-space normalisation technique that boosts out-of-distribution performance by +5.42% mAP for convolutional and +3.75% mAP for transformer-based models. Finally, we provide an in-depth analysis on the role of backgrounds in out-of-distribution behaviour recognition, including the so far unexplored impact of background durations (i.e., the count of background frames within foreground videos).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。