arXiv:2510.09203cs.CV2025-10被引 3

用视觉语言对齐方法提升牛行为识别,尤其在数据少时表现更好。

Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video

  • 将牛行为识别转为跨模态语义对齐,引入时间整合模块处理视频数据
  • 在6种行为上达到96.1%准确率,少样本下仍保持良好泛化能力
  • 构建了含1905段标注视频的CattleBehaviours6数据集,支持未来研究

真实农场环境中牲畜行为识别仍面临挑战,主要源于高质量标注视频数据集稀缺及大规模预训练数据与农业监控视频之间的领域差异。为此,我们提出Cattle-CLIP,一种面向农业场景的视觉-语言域自适应框架,将牛行为识别重构为跨模态语义对齐任务,而非单纯视觉分类。该框架通过引入时间整合模块,将图像级对比预训练拓展至基于视频的行为理解,实现跨时间的一致语义对齐。为缓解预训练模型所用网络级图文数据与真实牛只监控视频间的分布偏移,我们设计了定制化的增强策略和专用行为提示词。此外,我们构建了CattleBehaviours6数据集,包含1905个标注片段,涵盖六类室内行为,用于模型训练与评估。该数据集提供标准化行为定义,可作为未来研究的实用资源。在全监督与少样本学习场景下进行评估,重点考察数据稀缺条件下的行为识别性能。实验表明,Cattle-CLIP在六类行为上整体准确率达96.1%,对采食、饮水和站立反刍行为的召回率近乎完美,并在少样本设置中展现出强鲁棒性。

原文摘要 · Abstract (English)

Robust behaviour recognition in real-world farm environments remains challenging due to several data-related limitations, including the scarcity of well-annotated livestock video datasets and the substantial domain gap between large-scale pre-training corpora and agricultural surveillance footage. To address these challenges, we propose Cattle-CLIP, a domain-adaptive vision-language framework that reformulates cattle behaviour recognition as cross-modal semantic alignment rather than purely visual classification. Instead of directly fine-tuning visual backbones, Cattle-CLIP incorporates a temporal integration module to extend image-level contrastive pre-training to video-based behaviour understanding, enabling consistent semantic alignment across time. To mitigate the distribution shift between web-scale image-text data used for the pre-trained model and real-world cattle surveillance footage, we further introduce tailored augmentation strategies and specialised behaviour prompts. Furthermore, we construct CattleBehaviours6, a curated and behaviour-consistent video dataset comprising 1905 annotated clips across six indoor behaviours to support model training and evaluation. Beyond serving as a benchmark for our proposed method, the dataset provides a standardised ethogram definition, offering a practical resource for future research in livestock behaviour analysis. Cattle-CLIP is evaluated under both fully-supervised and few-shot learning scenarios, with a particular focus on data-scarce behaviour recognition, an important yet under-explored goal in livestock monitoring. Experiments show that Cattle-CLIP achieves 96.1% overall accuracy across six behaviours in supervised settings, with near-perfect recall for feeding, drinking and standing-ruminating behaviours, and demonstrates robust generalisation with limited data in few-shot scenarios.

行为识别多模态农业AI少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。