用图文对齐提升农田杂草分割泛化能力,跨场景更准更省数据。
Vision-Language Semantic Grounding for Multi-Domain Crop-Weed Segmentation
- 图文双编码融合,用自然语言指导特征优化,保留细节定位。
- 跨域训练下平均Dice达91.64%,最难点杂草识别提升15.42%。
- 少样本时仍稳定表现,适合真实农田部署的轻量高效模型。
细粒度作物-杂草分割对精准农业中的靶向除草至关重要。然而,现有深度学习模型因依赖特定数据集的视觉特征,在异质农业环境中泛化能力差。我们提出视觉-语言杂草分割(VL-WS)框架,通过语义对齐的领域无关表示实现像素级分割。该架构采用双编码器设计,冻结的CLIP图像文本嵌入与任务特定空间特征经由基于自然语言描述的FiLM层融合调制,使图像级文本描述引导通道级特征优化,同时保持精细空间定位。不同于仅在单一数据集上训练评估的先验方法,VL-WS在统一语料库上训练,涵盖近距地面影像(机器人平台)与高空无人机影像,覆盖多种作物、杂草种类、生长阶段及传感条件。在四个基准数据集上的实验表明,该框架有效,平均Dice得分达91.64%,较CNN基线提升4.98%;最困难杂草类别的提升最大,达到80.45%(基线65.03%),提升15.42%。此外,该模型在目标域监督有限时仍保持稳定性能,体现更强泛化性和数据效率。这些结果凸显视觉-语言对齐在可扩展、低标注成本分割模型中的潜力,适用于多样化真实农业场景。
原文摘要 · Abstract (English)
Fine-grained crop-weed segmentation is essential for enabling targeted herbicide application in precision agriculture. However, existing deep learning models struggle to generalize across heterogeneous agricultural environments due to reliance on dataset-specific visual features. We propose Vision-Language Weed Segmentation (VL-WS), a novel framework that addresses this limitation by grounding pixel-level segmentation in semantically aligned, domain-invariant representations. Our architecture employs a dual-encoder design, where frozen Contrastive Language-Image Pretraining (CLIP) embeddings and task-specific spatial features are fused and modulated via Feature-wise Linear Modulation (FiLM) layers conditioned on natural language captions. This design enables image level textual descriptions to guide channel-wise feature refinement while preserving fine-grained spatial localization. Unlike prior works restricted to training and evaluation on single-source datasets, VL-WS is trained on a unified corpus that includes close-range ground imagery (robotic platforms) and high-altitude UAV imagery, covering diverse crop types, weed species, growth stages, and sensing conditions. Experimental results across four benchmark datasets demonstrate the effectiveness of our framework, with VL-WS achieving a mean Dice score of 91.64% and outperforming the CNN baseline by 4.98%. The largest gains occur on the most challenging weed class, where VL-WS attains 80.45% Dice score compared to 65.03% for the best baseline, representing a 15.42% improvement. VL-WS further maintains stable weed segmentation performance under limited target-domain supervision, indicating improved generalization and data efficiency. These findings highlight the potential of vision-language alignment to enable scalable, label-efficient segmentation models deployable across diverse real-world agricultural domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。