用图模型自动找动物罕见行为,大幅降低标注成本
Sifting through the haystack -- efficiently finding rare animal behaviors in large-scale datasets
- 基于图的异常检测,从无标签姿态数据中识别罕见行为
- 仅需0.02%数据量即构建有效训练集,性能提升70%
- 适合研究稀有动物行为的科研人员,节省一半标注时间
在动物行为研究中,长期连续视频积累成大规模数据集,但感兴趣的罕见行为远少于常规行为,导致人工标注成本高昂。本文提出一种高效采样流程,仅需无标签的动物姿态或加速度数据作为输入,无需预设罕见行为的类型、数量或特征。方法基于近期的人类行为图异常检测模型,利用异常得分自动标记正常样本,将人工标注资源集中于异常区域。在实验室与野外采集的三个自由活动动物数据集上测试显示,该方法在运动类行为分析中表现优异,仅用少量标注预算即可取得良好效果。相比传统随机采样,平均性能提升70%,即使目标行为仅占数据总量0.02%,仍能有效构建分类器;当行为不罕见时,也至少减少50%标注工作量。
原文摘要 · Abstract (English)
In the study of animal behavior, researchers often record long continuous videos, accumulating into large-scale datasets. However, the behaviors of interest are often rare compared to routine behaviors. This incurs a heavy cost on manual annotation, forcing users to sift through many samples before finding their needles. We propose a pipeline to efficiently sample rare behaviors from large datasets, enabling the creation of training datasets for rare behavior classifiers. Our method only needs an unlabeled animal pose or acceleration dataset as input and makes no assumptions regarding the type, number, or characteristics of the rare behaviors. Our pipeline is based on a recent graph-based anomaly detection model for human behavior, which we apply to this new data domain. It leverages anomaly scores to automatically label normal samples while directing human annotation efforts toward anomalies. In research data, anomalies may come from many different sources (e.g., signal noise versus true rare instances). Hence, the entire labeling budget is focused on the abnormal classes, letting the user review and label samples according to their needs. We tested our approach on three datasets of freely-moving animals, acquired in the laboratory and the field. We found that graph-based models are particularly useful when studying motion-based behaviors in animals, yielding good results while using a small labeling budget. Our method consistently outperformed traditional random sampling, offering an average improvement of 70% in performance and creating datasets even when the behavior of interest was only 0.02% of the data. Even when the performance gain was minor (e.g., when the behavior is not rare), our method still reduced the annotation effort by half.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。