首个农业场景多标签实例级视觉定位数据集,解决田间作物与杂草定位难题。
Multi-label Instance-level Generalised Visual Grounding in Agriculture
- 构建gRef-CW数据集,支持农业场景下多标签、负向表达的视觉定位。
- 现有模型在田间条件下定位准确率低,暴露领域差距。
- 提出Weed-VG框架,融合层次化相关性评分与插值回归,提升定位精度。
理解农田图像,如检测植物并区分作物与杂草的个体实例,是精准农业的核心挑战。尽管视觉语言任务如图像描述和视觉问答取得进展,视觉定位(Visual Grounding, VG)在农业中仍未被探索。主要原因在于缺乏适合田间条件评估的基准数据集:植物外观高度相似、尺度多样,且目标可能不在图像中。为此,我们提出gRef-CW,首个面向农业通用视觉定位的数据集,包含负向表达。在gRef-CW上对当前先进定位模型进行基准测试,揭示显著领域差距,表明现有方法难以定位作物与杂草实例。基于此,我们提出Weed-VG,一种模块化框架,结合多标签层次相关性评分与插值驱动回归。Weed-VG实现了实例级视觉定位的进展,并为精准农业中视觉定位方法的发展提供了清晰基线。代码将在论文接受后发布。
原文摘要 · Abstract (English)
Understanding field imagery such as detecting plants and distinguishing individual crop and weed instances is a central challenge in precision agriculture. Despite progress in vision-language tasks like captioning and visual question answering, Visual Grounding (VG), localising language-referred objects, remains unexplored in agriculture. A key reason is the lack of suitable benchmark datasets for evaluating grounding models in field conditions, where many plants look highly similar, appear at multiple scales, and the referred target may be absent from the image. To address these limitations, we introduce gRef-CW, the first dataset designed for generalised visual grounding in agriculture, including negative expressions. Benchmarking current state-of-the-art grounding models on gRef-CW reveals a substantial domain gap, highlighting their inability to ground instances of crops and weeds. Motivated by these findings, we introduce Weed-VG, a modular framework that incorporates multi-label hierarchical relevance scoring and interpolation-driven regression. Weed-VG advances instance-level visual grounding and provides a clear baseline for developing VG methods in precision agriculture. Code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。