arXiv:2509.06427cs.CV2025-09被引 2

用语言模型零样本检测牛鼻子,无需标注数据

When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection

  • 用语言提示引导视觉模型,实现无标注的牛鼻定位
  • 在未见品种和环境下达到76.8% [email protected]
  • 适合畜牧监控中快速部署与跨环境应用

鼻部图案是牛识别中最有效的生物特征之一。快速准确地检测鼻部区域作为感兴趣区域,对自动视觉牛识别至关重要。早期方法依赖人工标注,耗时且不一致。近期基于监督模型(如YOLO)的自动化方法虽有效,但需大量标注数据,且训练依赖特定数据,难以泛化到新品种或未知场景。本研究提出一种基于Grounding DINO的零样本鼻部检测框架,该视觉-语言模型无需任务特训或标注数据,仅通过自然语言提示即可引导检测。该方法可实现多样品种与复杂环境下的可扩展、灵活定位。模型在标准测试上达到76.8%的[email protected],证明其无需标注数据仍具良好性能。据我们所知,这是首个面向实际产业场景、真正无标注的牛鼻检测解决方案。该框架为监督方法提供实用替代,有望提升牲畜监测中的适应性与部署便利性。

原文摘要 · Abstract (English)

Muzzle patterns are among the most effective biometric traits for cattle identification. Fast and accurate detection of the muzzle region as the region of interest is critical to automatic visual cattle identification.. Earlier approaches relied on manual detection, which is labor-intensive and inconsistent. Recently, automated methods using supervised models like YOLO have become popular for muzzle detection. Although effective, these methods require extensive annotated datasets and tend to be trained data-dependent, limiting their performance on new or unseen cattle. To address these limitations, this study proposes a zero-shot muzzle detection framework based on Grounding DINO, a vision-language model capable of detecting muzzles without any task-specific training or annotated data. This approach leverages natural language prompts to guide detection, enabling scalable and flexible muzzle localization across diverse breeds and environments. Our model achieves a mean Average Precision (mAP)@0.5 of 76.8\%, demonstrating promising performance without requiring annotated data. To our knowledge, this is the first research to provide a real-world, industry-oriented, and annotation-free solution for cattle muzzle detection. The framework offers a practical alternative to supervised methods, promising improved adaptability and ease of deployment in livestock monitoring applications.

牛识别零样本检测视觉语言模型畜牧监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。