arXiv:2604.03426cs.CV2026-04

用大模型+轻量逻辑实现猪群全天候自动追踪,减少标注依赖。

Automated Segmentation and Tracking of Group Housed Pigs Using Foundation Models

  • 用预训练视觉语言模型做通用视觉基础,再加模块化后处理适配猪场
  • 99%跟踪准确率,132分钟视频无身份切换,87%区域与轮廓匹配度
  • 适合想省数据、做长期监控的智能养殖系统开发者

基础模型(FM)正重塑计算机视觉,减少对任务特定监督学习的依赖,并利用大规模学习的通用视觉表征。在精准畜牧中,多数流程仍依赖需大量标注数据、重复训练和农场调优的监督模型。本研究提出以基础模型为中心的猪群自动监测工作流:预训练视觉-语言模型作为通用视觉主干,通过模块化后处理实现农场特异性适配。首先在1,418张标注图像上应用Grounding-DINO建立检测基线,白天检测准确率高,但夜间和严重遮挡下性能下降,因此引入时序追踪逻辑。基于这些检测结果,在550段一分钟视频上评估了Grounded-SAM2的短时视频分割;经后处理后,4,927条活跃轨迹中超过80%完全正确,主要错误源于掩码不准或标签重复。为支持长时间身份一致性,进一步构建集成初始化、追踪、匹配、掩码优化、重识别与事后质量控制的长时追踪流水线。该系统在连续132分钟视频上表现稳定,132个均匀采样真值帧中,平均区域相似度(J)为0.83,轮廓精度(F)为0.92,J&F为0.87,MOTA为0.99,MOTP为90.7%,且无身份切换。整体表明,基础模型先验知识结合轻量任务逻辑,可实现可扩展、低标注、长周期的养猪监控。

原文摘要 · Abstract (English)

Foundation models (FM) are reshaping computer vision by reducing reliance on task-specific supervised learning and leveraging general visual representations learned at scale. In precision livestock farming, most pipelines remain dominated by supervised learning models that require extensive labeled data, repeated retraining, and farm-specific tuning. This study presents an FM-centered workflow for automated monitoring of group-housed nursery pigs, in which pretrained vision-language FM serve as general visual backbones and farm-specific adaptation is achieved through modular post-processing. Grounding-DINO was first applied to 1,418 annotated images to establish a baseline detection performance. While detection accuracy was high under daytime conditions, performance degraded under night-vision and heavy occlusion, motivating the integration of temporal tracking logic. Building on these detections, short-term video segmentation with Grounded-SAM2 was evaluated on 550 one-minute video clips; after post-processing, over 80% of 4,927 active tracks were fully correct, with most remaining errors arising from inaccurate masks or duplicated labels. To support identity consistency over an extended time, we further developed a long-term tracking pipeline integrating initialization, tracking, matching, mask refinement, re-identification, and post-hoc quality control. This system was evaluated on a continuous 132-minute video and maintained stable identities throughout. On 132 uniformly sampled ground-truth frames, the system achieved a mean region similarity (J) of 0.83, contour accuracy (F) of 0.92, J&F of 0.87, MOTA of 0.99, and MOTP of 90.7%, with no identity switches. Overall, this work demonstrates how FM prior knowledge can be combined with lightweight, task-specific logic to enable scalable, label-efficient, and long-duration monitoring in pig production.

猪群追踪基础模型视频分割智能养殖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。