arXiv:2505.00569cs.CV2025-05中稿 · the poster session…被引 4

将光流与视频帧融合,提升动物行为识别精度

AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis

  • 用光流与帧交替输入,增强运动信息捕捉
  • 在Animal Kingdom数据集上超越现有最佳方法
  • 适合需要精细动作识别的动物行为研究

近年来,深度学习在动物行为识别中受到广泛关注,尤其是利用CLIP等预训练视觉语言模型,因其在多种下游任务中表现出色。然而,将其应用于动物行为识别面临两大挑战:如何融入运动信息,以及设计有效的时序建模方案。本文提出AnimalMotionCLIP,通过在CLIP框架中交错插入视频帧与光流信息来解决这些问题。同时,提出了三种基于分类器聚合的时序建模方法:密集、半密集和稀疏,并进行比较。实验表明,该方法能准确识别精细时序行为,在Animal Kingdom数据集上的表现优于现有先进方法。

原文摘要 · Abstract (English)

Recently, there has been a surge of interest in applying deep learning techniques to animal behavior recognition, particularly leveraging pre-trained visual language models, such as CLIP, due to their remarkable generalization capacity across various downstream tasks. However, adapting these models to the specific domain of animal behavior recognition presents two significant challenges: integrating motion information and devising an effective temporal modeling scheme. In this paper, we propose AnimalMotionCLIP to address these challenges by interleaving video frames and optical flow information in the CLIP framework. Additionally, several temporal modeling schemes using an aggregation of classifiers are proposed and compared: dense, semi dense, and sparse. As a result, fine temporal actions can be correctly recognized, which is of vital importance in animal behavior analysis. Experiments on the Animal Kingdom dataset demonstrate that AnimalMotionCLIP achieves superior performance compared to state-of-the-art approaches.

动物行为CLIP时序建模光流融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。